🤖 AI Summary
This study addresses the insufficient modeling of embedding-based semantic structures of assessment items and the limited interpretability of context scores in psychometrics. To overcome these limitations, it proposes constructing context scores using external corpus similarity and introduces a partially specified two-step factor analysis procedure, with technical validation integrating vector embeddings, the Bayesian Information Criterion (BIC), and Q-matrix diagnostic models. Applied to TIMSS mathematics items, this work reveals a stable seven-cluster semantic structure that distinguishes multiple association types. Furthermore, it demonstrates that conditional semantic representations outperform single-factor models, offering new directions for text-assisted response calibration.
📝 Abstract
Contextual scores represent assessment items through their similarities to reference words in an external corpus. We examine the semantic structure of scores for 40 TIMSS mathematics scored units using a partially specified two-step factor procedure. A search across factor counts identifies a persistent seven-group structure under the featured construction. Subsequent comparisons consistently favor a general dimension alongside group associations, although individual group memberships remain sensitive to some specification choices. Item examples distinguish recurring, cross-domain, sensitive, and imposed associations. Simpler and unrestricted references clarify the contribution and limits of the anchored representation: it improves on a single factor but does not achieve the lowest working Bayesian information criterion (BIC). A separate response benchmark compares three initial Q constructions and their Hull-PVAF revisions under higher-order and saturated attribute distributions. Among these diagnostic models, BIC favors the official content framework and the Akaike information criterion (AIC) favors its direct four-factor augmentation, but a matched unidimensional two-parameter logistic model has lower AIC and BIC than all twelve conditions. These findings support a conditional semantic representation while limiting direct diagnostic interpretation. We discuss learned text-assisted response calibration as a prospective application requiring a larger calibrated item bank and independent evaluation.