🤖 AI Summary
A longstanding issue in IRT simulation—“reliability omission”—treats reliability as an implicit byproduct rather than an explicit, controllable design parameter, resulting in ambiguous signal-to-noise ratios. This paper formally defines the IRT inverse-design problem and introduces the first simulation framework enabling precise, user-specified control of marginal reliability—elevating it from an output metric to an explicit input parameter. We innovatively distinguish and calibrate two reliability types—equivalent-class (EQC) and stochastic-approximation (SAC)—yielding two deterministic and stochastic algorithms, respectively: EQC achieves near-exact calibration, while SAC ensures unbiased estimation under non-normal latent traits and realistic item pools. Leveraging Jensen’s inequality for theoretical analysis, we validate the framework across 960 experimental conditions. We publicly release the R package *IRTsimrel*, enabling standardized, reliability-aware IRT simulation.
📝 Abstract
Monte Carlo simulations are the primary methodology for evaluating Item Response Theory (IRT) methods, yet marginal reliability - the fundamental metric of data informativeness - is rarely treated as an explicit design factor. Unlike in multilevel modeling where the intraclass correlation (ICC) is routinely manipulated, IRT studies typically treat reliability as an incidental outcome, creating a "reliability omission" that obscures the signal-to-noise ratio of generated data. To address this gap, we introduce a principled framework for reliability-targeted simulation, transforming reliability from an implicit by-product into a precise input parameter. We formalize the inverse design problem, solving for a global discrimination scaling factor that uniquely achieves a pre-specified target reliability. Two complementary algorithms are proposed: Empirical Quadrature Calibration (EQC) for rapid, deterministic precision, and Stochastic Approximation Calibration (SAC) for rigorous stochastic estimation. A comprehensive validation study across 960 conditions demonstrates that EQC achieves essentially exact calibration, while SAC remains unbiased across non-normal latent distributions and empirical item pools. Furthermore, we clarify the theoretical distinction between average-information and error-variance-based reliability metrics, showing they require different calibration scales due to Jensen's inequality. An accompanying open-source R package, IRTsimrel, enables researchers to standardize reliability as a controlled experimental input.