Score
Designs and implements procedures to estimate and validate Classical Test Theory item parameters (e.g., item difficulty and item discrimination) from response data, including algorithms that compute discrimination from observed responses and aggregate estimates across respondent pools. Builds and evaluates response-based calibration workflows—often using synthetic respondent generators—to perform calibration, compare synthetic and human-derived parameter estimates, and analyze the robustness of discrimination estimates.
This study addresses the potential pitfalls of directly applying item response theory (IRT)—originally designed for human assessment—to the evaluation of artificial intelligence systems, where mismatched data-generating mechanisms may compromise inference validity. It presents the first systematic evaluation of IRT’s applicability to large language model benchmarks, examining the feasibility, scalability, and reliability of four estimation approaches—marginal maximum likelihood, Markov chain Monte Carlo (MCMC), variational inference, and neural pseudo-twin estimators—across 18,000 simulated conditions. The findings reveal that classical methods are computationally infeasible at scale, while scalable alternatives introduce bias when the number of models is small or their ability distribution deviates from normality. The work quantifies, for the first time, the failure boundaries of IRT in AI evaluation and establishes required sample sizes and diagnostic criteria for its reliable application.
This study addresses the high cost and time demands of traditional item response theory (IRT) parameter calibration, which relies on real student response data. The authors propose a novel approach leveraging large language models: by fine-tuning the Qwen-3 dense model with LoRA, they generate synthetic responses conditioned on discrete ability descriptors to simulate students at varying proficiency levels. This enables reconstruction of item characteristic curves (ICCs) and estimation of IRT parameters without any real-world response data. To the best of the authors’ knowledge, this is the first method to implicitly model psychometric properties in a data-free setting. The approach demonstrates particularly strong performance in estimating item discrimination and achieves competitive or superior results compared to existing baselines on both sixth-grade English Language Arts items and the BEA 2024 dataset, confirming its effectiveness and practical utility.
A longstanding issue in IRT simulation—“reliability omission”—treats reliability as an implicit byproduct rather than an explicit, controllable design parameter, resulting in ambiguous signal-to-noise ratios. This paper formally defines the IRT inverse-design problem and introduces the first simulation framework enabling precise, user-specified control of marginal reliability—elevating it from an output metric to an explicit input parameter. We innovatively distinguish and calibrate two reliability types—equivalent-class (EQC) and stochastic-approximation (SAC)—yielding two deterministic and stochastic algorithms, respectively: EQC achieves near-exact calibration, while SAC ensures unbiased estimation under non-normal latent traits and realistic item pools. Leveraging Jensen’s inequality for theoretical analysis, we validate the framework across 960 experimental conditions. We publicly release the R package *IRTsimrel*, enabling standardized, reliability-aware IRT simulation.
This study addresses the challenge of item parameter recovery in the Multidimensional Graded Response Model (MGRM). Methodologically, it implements a fully reproducible R framework that integrates Monte Carlo simulation for generating multidimensional ordinal response data, maximum likelihood estimation for fitting the three-dimensional GRM, and systematic evaluation of estimation accuracy via bias and root mean square error (RMSE). Robustness is assessed across manipulated conditions: test length (20 vs. 40 items), interdimensional correlation (0.3 vs. 0.7), and sample size (N = 2000). The key contribution is the first comprehensive, end-to-end R workflow—encompassing data generation, parameter estimation, diagnostic evaluation, and visualization (using ggplot2)—designed for both pedagogical clarity and cross-disciplinary methodological transfer. This framework substantially lowers the technical barrier for researchers without formal psychometric training to apply MGRM rigorously.
This study addresses the challenge of inaccurate item parameter estimation for new items in computerized adaptive testing, which arises from sparse response data and degrades scoring quality. To tackle this issue, the authors propose a content-driven compact calibration method that leverages a pretrained shallow ReLU network to map handcrafted item features into a low-dimensional embedding space. This embedding is then decoupled and integrated with a linear, interpretable three-parameter logistic (3PL) item response theory (IRT) model, eliminating the need for online parameter tuning during deployment. The approach achieves, for the first time, a disentangled joint optimization of neural embeddings and interpretable IRT parameters. Evaluated on two task types in the Duolingo English Test, the method demonstrates superior performance using only a six-dimensional embedding, outperforming larger models and confirming its efficiency and effectiveness.
This study presents the first systematic evaluation of whether large language models (LLMs) can capture item discriminability—the core psychometric property that enables reading comprehension items to differentiate between students of varying ability levels. Under a zero-shot setting, 42 LLMs were assessed using two complementary approaches: direct prediction of item discriminability and classical test theory (CTT) calibration based on model-synthesized responses. Results indicate that although LLM outputs contain non-random signals, their discriminability patterns show limited alignment with human-calibrated benchmarks: the best Spearman correlation for direct prediction reaches only 0.152, improving modestly to 0.241 with the CTT approach—still far from reliably replicating the discrimination structure inherent in human responses. This work thus reveals significant limitations of current LLMs in psychometric modeling.
This study addresses the high cost of calibrating item parameters in multidimensional item response theory (MIRT), which typically requires extensive respondent data. To circumvent traditional calibration, the authors propose leveraging pretrained text embeddings to directly recover directional loadings from item stems and demonstrate their efficacy within multidimensional computerized adaptive testing. They introduce a novel geometric diagnostic based on the condition number of the loading matrix to assess item bank suitability and provide the first systematic analysis of the accuracy and uncertainty of semantic embeddings in MIRT. Experimental results show that embedding-derived latent trait profiles achieve a correlation of 0.825 (0.857 with fitted parameters), significantly outperforming a lexical overlap baseline (0.752). However, posterior variance inflates nearly fourfold due to high collinearity among embedding dimensions, evidenced by an average cosine similarity of 0.90 and a condition number of 137.
This study addresses the challenge of scale interference in operational latent variables during embedded pilot item calibration under few-shot adaptive routing. To mitigate this issue, the authors propose a Bayesian modular framework that severs feedback from pilot responses to operational latent variables, instead calibrating each pilot item independently via local logistic regression using fixed predictors derived from the posterior of operational traits. The approach innovatively incorporates an analytical debiasing mapping, enabling unbiased recovery of item parameters under a normal-ogive approximation. By integrating Firth’s penalized likelihood, multivariate delta-method covariance estimation, and Rubin’s pooling rules, the method achieves marginal likelihood calibration without per-item numerical integration while propagating parameter uncertainty into standard errors. Simulation results demonstrate substantially reduced attenuation bias and yield confidence interval coverage ranging from nominal to conservative under constrained conditions such as missing-at-random routing.
This study addresses the cold-start problem in item parameter estimation when newly developed test items lack empirical response data. The authors propose a prediction approach leveraging textual embeddings and regularized regression, accompanied by an evaluation framework integrating resampling-based cross-validation, reliability ceilings, and design ceilings. Innovatively employing a dual “ceiling” analysis, the work demonstrates that differences in parameter predictability stem primarily from measurement reliability rather than the strength of textual information, underscoring the necessity of repeated validation. In the EEDI mathematics item bank, predicted difficulty parameters achieved an R² of 0.53, representing 57% of the reliability ceiling, whereas pseudo-guessing parameters in the three-parameter logistic model proved largely unpredictable due to near-zero reliability ceilings. BEA benchmark experiments further reveal that relying solely on RMSE can obscure extremely low explained variance, highlighting the critical role of dimensionless metrics in model evaluation.