Score
Designs and fits item response theory models in which items or conditions are described by continuous difficulty/effect functions and observed outcomes are generated by a probabilistic response function (e.g., logistic). Builds estimation and inference procedures that recover individuals' latent abilities and continuous item-difficulty curves, including methods to impose expert priors or anchors and to recover parametric or nonparametric response relationships.
Social scientists frequently infer dynamic latent variables—such as public opinion or ideological positioning—from longitudinal ordinal indicators, yet conventional Item Response Theory (IRT) models lack flexibility in characterizing continuous-time latent evolution and nonlinear item response functions. To address this, we propose the Generalized Dynamic Gaussian Process IRT (GPIRT) framework, the first to integrate Bayesian nonparametric IRT with Gaussian process (GP) temporal modeling. GPIRT employs dynamic GP priors to represent latent trajectories continuously over time, supports data-driven estimation of diverse nonlinear response functions, and introduces a tailored MCMC algorithm for scalable posterior inference. Simulation studies demonstrate substantial gains in accuracy and robustness over state-of-the-art dynamic IRT methods. Empirical applications to economic sentiment and congressional abortion ideology uncover fine-grained, time-varying patterns that standard approaches fail to detect.
Four-parameter item response theory (4P-IRT) lacks a corresponding factor analysis (FA) counterpart, hindering unified latent variable modeling and systematic correction of response biases—namely, guessing and inattention—for binary items. Method: This paper introduces the four-parameter factor analysis (4P-FA) model, establishing its analytical equivalence to 4P-IRT for the first time. We develop an identifiable Bayesian inference framework that jointly estimates all four item parameters and bias-corrected person-specific latent traits, implemented efficiently in R/Python. Contribution/Results: Empirical applications to real-world college admission and anxiety assessment data demonstrate that 4P-FA accurately disentangles and corrects for guessing and inattention effects, substantially improving the validity and interpretability of latent trait estimation. By bridging the theoretical gap between IRT and FA, this work provides a novel paradigm for response-bias modeling in psychometrics.
This study addresses the challenge of item parameter recovery in the Multidimensional Graded Response Model (MGRM). Methodologically, it implements a fully reproducible R framework that integrates Monte Carlo simulation for generating multidimensional ordinal response data, maximum likelihood estimation for fitting the three-dimensional GRM, and systematic evaluation of estimation accuracy via bias and root mean square error (RMSE). Robustness is assessed across manipulated conditions: test length (20 vs. 40 items), interdimensional correlation (0.3 vs. 0.7), and sample size (N = 2000). The key contribution is the first comprehensive, end-to-end R workflow—encompassing data generation, parameter estimation, diagnostic evaluation, and visualization (using ggplot2)—designed for both pedagogical clarity and cross-disciplinary methodological transfer. This framework substantially lowers the technical barrier for researchers without formal psychometric training to apply MGRM rigorously.
This study investigates the bias in ability estimation and inflation of parameter uncertainty arising from misspecifying a noncompensatory multidimensional item response theory (MIRT) model as compensatory. Using asymptotic variance analysis and theoretical derivation, we establish— for the first time—that ability estimates are systematically overestimated near the origin, and characterize the bias as bimodal: overestimation for low-ability examinees and underestimation for high-ability examinees. We derive the distributional properties of the estimation bias under misspecification and obtain a theoretical upper bound on the asymptotic variance of ability estimates. Results demonstrate that model misspecification not only distorts point estimates but also substantially amplifies estimation uncertainty—particularly when latent dimensions exhibit strong interaction. Our findings provide quantifiable theoretical criteria and diagnostic tools for MIRT model selection, thereby enhancing the robustness and scientific rigor of multidimensional test modeling.
To address the challenges of modeling student ability dynamics and ensuring interpretability in sparse longitudinal educational assessment data, this paper proposes the Dynamic Bayesian Item Response Model (D-BIRD). D-BIRD innovatively decomposes student ability into two separable, interpretable latent components: population-level learning trends and individual-specific deviations, jointly modeled via a hierarchical Bayesian framework that captures their temporal evolution. Methodologically, it integrates a dynamic factor structure with variational inference and MCMC-based posterior estimation to enable efficient, scalable inference. Simulation studies demonstrate high parameter recovery accuracy. On real-world personalized learning data, D-BIRD achieves a 12.7% reduction in MAE for ability trajectory prediction compared to baselines, while enabling fine-grained identification of population-level learning patterns and individual developmental attribution. This provides an interpretable, principled foundation for intelligent educational diagnosis and adaptive intervention.
Existing online learning analytics struggle to capture the nonlinear dynamics of student ability, often requiring a prespecified number of clusters and failing to effectively model the relationship between engagement behaviors and ability evolution. This work proposes a Bayesian nonparametric dynamic item response theory framework that employs B-spline basis functions to flexibly characterize the nonlinear influence of participation on ability drift. By incorporating a Mixture-of-Finite-Mixtures prior, the model automatically infers the number of latent learner subgroups, enabling unsupervised clustering and longitudinal tracking of individual ability trajectories. Applied to data from 198 undergraduate students in a statistics course, the model identified four distinct learner types—struggling-declining, low-stable, mainstream-stable, and high-improving—revealing highly stable ability trajectories and no significant predictive effect of participation volume on ability drift.
This study systematically investigates the prevalence and psychometric consequences of deviations from the normality assumption of latent trait distributions in item response theory (IRT). Analyzing 504 real-world datasets, the authors employed flexible nonparametric and semiparametric methods to estimate trait distributions and compared these against conventional normality assumptions, evaluating impacts on reliability, item parameters, predicted responses, and individual scores. The large-scale empirical analysis reveals—for the first time—that more than half of the datasets exhibit cumulative distribution discrepancies exceeding 10 percentage points (with roughly one-fifth surpassing 20 points), frequently manifesting as skewed, heavy-tailed, flat, or multimodal shapes, with substantial variation across domains. These deviations exert particularly pronounced effects when full models are refitted. The findings underscore the necessity of routinely reporting distributional sensitivity analyses in IRT applications.
This study proposes a novel approach leveraging the multimodal large language model Qwen-VL 3.5 to implicitly learn item response theory (IRT) parameterized response curves directly from multiple-choice questions containing both text and images, without requiring explicit parameter fitting. Through supervised fine-tuning and prompt engineering, the model reproduces option-level response probabilities conditioned on student ability, simultaneously modeling both the three-parameter logistic (3PL) model and the multiple-choice model (MCM). Experimental results demonstrate that the method accurately approximates true item difficulty parameters on held-out test sets and effectively captures systematic error patterns across students of varying abilities. To the best of our knowledge, this work represents the first successful end-to-end reconstruction of IRT response functions using a multimodal large language model.
This study addresses the conflation of environmental difficulty and individual skill in traditional assessments of outdoor activity suitability, which typically rely on a single expert-derived curve. The authors introduce continuous item response theory to this domain for the first time, jointly modeling rider ability and trail difficulty using rider performance, terrain conditions, and binary outcomes. Success probability is characterized via a sigmoidal function, with a physics-informed expert curve serving as a prior for difficulty. Parameters are estimated through marginal maximum likelihood using Gaussian–Hermite quadrature, and graph connectivity constraints ensure identifiability. Experiments demonstrate that the model achieves a skill recovery correlation of 0.96 on synthetic data, locates minimum difficulty with an error under three units, and improves Brier skill scores by 0.33 over the expert-curve baseline, effectively disentangling and quantifying intrinsic difficulty from individual capability.
Traditional multidimensional item response theory is constrained by the assumption that latent traits follow a Gaussian distribution, which often fails to capture complex structures such as skewness, heavy tails, or multimodality, leading to biased parameter estimates. This work proposes the first integration of normalizing flows into this framework, leveraging invertible neural networks to model latent traits as flexible transformations of a simple base distribution. By combining conditional flows with variational inference, the approach jointly learns item parameters, the latent trait distribution, and its posterior. Simulation studies demonstrate that the method substantially improves the accuracy of both parameter and trait recovery under non-normal conditions. Furthermore, application to real-world personality data confirms its capacity to effectively model intricate latent distributions.