Score
Designs and fits probabilistic measurement models (Rasch and broader item response theory variants) that map observed item or response data to one or more latent proficiency traits, estimating item parameters (e.g., difficulty, discrimination, guessing) and person ability on a common scale while performing calibration, scaling/equating, and model-fit diagnostics. Produces calibrated item characteristic functions, information curves, and continuous latent proficiency scores and validated item parameters to be used as features or inputs for downstream models and adaptive assessment systems.
Estimating item parameters in the Rasch model under large-scale sparse observational data poses a fundamental challenge—balancing minimax optimality, uncertainty quantification, and computational feasibility. Method: We propose Randomized Pairwise Maximum Likelihood Estimation (RP-MLE) and its multi-replicate variant (MRP-MLE), which construct randomized pairwise comparisons among items to achieve dimensionality reduction while preserving sample independence. Contribution/Results: We establish that RP-MLE achieves finite-sample ℓ∞-norm minimax optimality and attains the information-theoretic lower bound asymptotically. Crucially, it provides the first rigorous asymptotic distributional characterization—and thus verifiable confidence intervals—for item parameter estimates under sparsity. Extensive simulations and real-data experiments demonstrate that RP-MLE and MRP-MLE significantly outperform classical estimators in both estimation accuracy and statistical inference, effectively breaking the traditional dependence of estimators on data density.
This study addresses the challenge of calibrating explanatory item response theory (IRT) parameters under large-scale sparse data—a common scenario in adaptive testing where examinees respond to only a small fraction of items. To this end, the authors propose a Bayesian multidimensional explanatory IRT model, complemented by a tailored MCMC algorithm, a sparsity-aware data structure, and a high-performance computing engine. This integrated approach enables, for the first time, efficient and stable calibration of IRT parameters in ultra-large-scale sparse psychometric datasets. The resulting scalable SPICE calibration engine supports diverse psychometric applications and demonstrates strong performance and practical utility in real-world contexts such as computerized adaptive testing and automated item bank generation.
This study addresses the high cost and time demands of traditional item response theory (IRT) parameter calibration, which relies on real student response data. The authors propose a novel approach leveraging large language models: by fine-tuning the Qwen-3 dense model with LoRA, they generate synthetic responses conditioned on discrete ability descriptors to simulate students at varying proficiency levels. This enables reconstruction of item characteristic curves (ICCs) and estimation of IRT parameters without any real-world response data. To the best of the authors’ knowledge, this is the first method to implicitly model psychometric properties in a data-free setting. The approach demonstrates particularly strong performance in estimating item discrimination and achieves competitive or superior results compared to existing baselines on both sixth-grade English Language Arts items and the BEA 2024 dataset, confirming its effectiveness and practical utility.
To address the computational complexity and convergence difficulties inherent in joint estimation of measurement and structural models in item response theory (IRT), this paper proposes a two-step maximum likelihood estimation procedure: first, estimating measurement model parameters independently; second, estimating structural model parameters with measurement parameters held fixed. This work provides the first systematic theoretical justification—under settings involving continuous latent variables and categorical observed variables—of the statistical consistency, robustness, and computational efficiency of the two-step approach. Compared to conventional one-step estimation (prone to non-convergence) and three-step methods (susceptible to bias accumulation), the proposed method offers conceptual clarity, implementation simplicity, reliable standard errors, and stable convergence. Extensive simulation studies and empirical analyses validate its efficacy and generalizability across diverse latent variable models. The framework establishes a novel, general-purpose, flexible, and practical estimation paradigm for educational measurement, psychometrics, and related fields.
This study addresses the challenge of inaccurate item parameter estimation for new items in computerized adaptive testing, which arises from sparse response data and degrades scoring quality. To tackle this issue, the authors propose a content-driven compact calibration method that leverages a pretrained shallow ReLU network to map handcrafted item features into a low-dimensional embedding space. This embedding is then decoupled and integrated with a linear, interpretable three-parameter logistic (3PL) item response theory (IRT) model, eliminating the need for online parameter tuning during deployment. The approach achieves, for the first time, a disentangled joint optimization of neural embeddings and interpretable IRT parameters. Evaluated on two task types in the Duolingo English Test, the method demonstrates superior performance using only a six-dimensional embedding, outperforming larger models and confirming its efficiency and effectiveness.
This study systematically investigates the prevalence and psychometric consequences of deviations from the normality assumption of latent trait distributions in item response theory (IRT). Analyzing 504 real-world datasets, the authors employed flexible nonparametric and semiparametric methods to estimate trait distributions and compared these against conventional normality assumptions, evaluating impacts on reliability, item parameters, predicted responses, and individual scores. The large-scale empirical analysis reveals—for the first time—that more than half of the datasets exhibit cumulative distribution discrepancies exceeding 10 percentage points (with roughly one-fifth surpassing 20 points), frequently manifesting as skewed, heavy-tailed, flat, or multimodal shapes, with substantial variation across domains. These deviations exert particularly pronounced effects when full models are refitted. The findings underscore the necessity of routinely reporting distributional sensitivity analyses in IRT applications.
This study addresses the challenge of scale interference in operational latent variables during embedded pilot item calibration under few-shot adaptive routing. To mitigate this issue, the authors propose a Bayesian modular framework that severs feedback from pilot responses to operational latent variables, instead calibrating each pilot item independently via local logistic regression using fixed predictors derived from the posterior of operational traits. The approach innovatively incorporates an analytical debiasing mapping, enabling unbiased recovery of item parameters under a normal-ogive approximation. By integrating Firth’s penalized likelihood, multivariate delta-method covariance estimation, and Rubin’s pooling rules, the method achieves marginal likelihood calibration without per-item numerical integration while propagating parameter uncertainty into standard errors. Simulation results demonstrate substantially reduced attenuation bias and yield confidence interval coverage ranging from nominal to conservative under constrained conditions such as missing-at-random routing.
This study addresses the challenge of determining the number of latent dimensions in multidimensional graded response models by proposing an adaptive Bayesian framework for dimensionality selection. The approach introduces a cumulative ordered spike-and-slab (COSS) prior on the column variances of the item loading matrix, which automatically shrinks redundant dimensions while preserving meaningful structure. Coupled with Albert–Chib latent variable augmentation, it enables an efficient Gibbs sampler. A key innovation lies in integrating ordered shrinkage with Bayesian nonparametric principles, allowing adaptive inference of the latent dimensionality without pre-specifying a candidate set and naturally quantifying uncertainty in dimension selection. Simulation studies and analyses of real psychological assessment data demonstrate that the method accurately recovers true dimensional structures, yields more precise parameter estimates, and maintains computational efficiency.
Traditional multidimensional item response theory is constrained by the assumption that latent traits follow a Gaussian distribution, which often fails to capture complex structures such as skewness, heavy tails, or multimodality, leading to biased parameter estimates. This work proposes the first integration of normalizing flows into this framework, leveraging invertible neural networks to model latent traits as flexible transformations of a simple base distribution. By combining conditional flows with variational inference, the approach jointly learns item parameters, the latent trait distribution, and its posterior. Simulation studies demonstrate that the method substantially improves the accuracy of both parameter and trait recovery under non-normal conditions. Furthermore, application to real-world personality data confirms its capacity to effectively model intricate latent distributions.