When Can Text Embeddings Replace Item Calibration? A Geometric Diagnostic for Semantic Loadings in Multidimensional Adaptive Testing

📅 2026-08-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high cost of calibrating item parameters in multidimensional item response theory (MIRT), which typically requires extensive respondent data. To circumvent traditional calibration, the authors propose leveraging pretrained text embeddings to directly recover directional loadings from item stems and demonstrate their efficacy within multidimensional computerized adaptive testing. They introduce a novel geometric diagnostic based on the condition number of the loading matrix to assess item bank suitability and provide the first systematic analysis of the accuracy and uncertainty of semantic embeddings in MIRT. Experimental results show that embedding-derived latent trait profiles achieve a correlation of 0.825 (0.857 with fitted parameters), significantly outperforming a lexical overlap baseline (0.752). However, posterior variance inflates nearly fourfold due to high collinearity among embedding dimensions, evidenced by an average cosine similarity of 0.90 and a condition number of 137.
📝 Abstract
Multidimensional item response theory relies on calibrated item parameters, such as discrimination and category threshold values, which are usually estimated from large samples of human test responses. This study investigates whether the directional loadings of these parameters can be recovered directly from item text using pre-trained sentence embeddings, avoiding the need for initial item calibration. Using the open-source IPIP Big-Five dataset ($n=19{,}719$; 50 items), we built a multidimensional computerized adaptive testing (CAT) simulation using D-optimal item selection. We compared three item loading sources: fitted graded response model parameters, semantic text embeddings, and a lexical baseline. In simulation, semantic embeddings recovered latent trait profiles almost as accurately as fitted parameters (correlation $0.825$ vs. $0.857$), performing noticeably better than simple word overlap ($0.752$). However, the embedding-based model produced inflated posterior variance, showing nearly four times higher measurement uncertainty despite accurate point estimates. We attribute this to collinearity across dimensions, as embedding-derived loadings pointed in similar directions across traits (condition number $137$ vs. $1.0$; mean trait cosine $0.90$). This outcome reflects the shared vocabulary common in personality items. We propose a simple diagnostic metric based on the loading matrix condition number to evaluate whether an item bank is suitable for text-derived loadings prior to testing.
Problem

Research questions and friction points this paper is trying to address.

multidimensional adaptive testing
item calibration
text embeddings
semantic loadings
item response theory
Innovation

Methods, ideas, or system contributions that make the work stand out.

text embeddings
multidimensional adaptive testing
item calibration
semantic loadings
condition number
🔎 Similar Papers
No similar papers found.
A
Amirreza Mehrabi
School of Engineering Education & Department of Electrical and Computer Engineering, Purdue University