🤖 AI Summary
This study addresses the high cost of calibrating item parameters in multidimensional item response theory (MIRT), which typically requires extensive respondent data. To circumvent traditional calibration, the authors propose leveraging pretrained text embeddings to directly recover directional loadings from item stems and demonstrate their efficacy within multidimensional computerized adaptive testing. They introduce a novel geometric diagnostic based on the condition number of the loading matrix to assess item bank suitability and provide the first systematic analysis of the accuracy and uncertainty of semantic embeddings in MIRT. Experimental results show that embedding-derived latent trait profiles achieve a correlation of 0.825 (0.857 with fitted parameters), significantly outperforming a lexical overlap baseline (0.752). However, posterior variance inflates nearly fourfold due to high collinearity among embedding dimensions, evidenced by an average cosine similarity of 0.90 and a condition number of 137.
📝 Abstract
Multidimensional item response theory relies on calibrated item parameters, such as discrimination and category threshold values, which are usually estimated from large samples of human test responses. This study investigates whether the directional loadings of these parameters can be recovered directly from item text using pre-trained sentence embeddings, avoiding the need for initial item calibration. Using the open-source IPIP Big-Five dataset ($n=19{,}719$; 50 items), we built a multidimensional computerized adaptive testing (CAT) simulation using D-optimal item selection. We compared three item loading sources: fitted graded response model parameters, semantic text embeddings, and a lexical baseline. In simulation, semantic embeddings recovered latent trait profiles almost as accurately as fitted parameters (correlation $0.825$ vs. $0.857$), performing noticeably better than simple word overlap ($0.752$). However, the embedding-based model produced inflated posterior variance, showing nearly four times higher measurement uncertainty despite accurate point estimates. We attribute this to collinearity across dimensions, as embedding-derived loadings pointed in similar directions across traits (condition number $137$ vs. $1.0$; mean trait cosine $0.90$). This outcome reflects the shared vocabulary common in personality items. We propose a simple diagnostic metric based on the loading matrix condition number to evaluate whether an item bank is suitable for text-derived loadings prior to testing.