Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models

πŸ“… 2026-01-05
πŸ›οΈ arXiv.org
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the high cost and time demands of traditional item response theory (IRT) parameter calibration, which relies on real student response data. The authors propose a novel approach leveraging large language models: by fine-tuning the Qwen-3 dense model with LoRA, they generate synthetic responses conditioned on discrete ability descriptors to simulate students at varying proficiency levels. This enables reconstruction of item characteristic curves (ICCs) and estimation of IRT parameters without any real-world response data. To the best of the authors’ knowledge, this is the first method to implicitly model psychometric properties in a data-free setting. The approach demonstrates particularly strong performance in estimating item discrimination and achieves competitive or superior results compared to existing baselines on both sixth-grade English Language Arts items and the BEA 2024 dataset, confirming its effectiveness and practical utility.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationNatural Language Processing: (Large) Language ModelsCognitive Modeling & Cognitive Systems: Simulating Human Behavior

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Large language models for search
πŸ“ Abstract
Traditional methods for determining assessment item parameters, such as difficulty and discrimination, rely heavily on expensive field testing to collect student performance data for Item Response Theory (IRT) calibration. This study introduces a novel approach that implicitly models these psychometric properties by fine-tuning Large Language Models (LLMs) to simulate student responses across a spectrum of latent abilities. Leveraging the Qwen-3 dense model series and Low-Rank Adaptation (LoRA), we train models to generate responses to multiple choice questions conditioned on discrete ability descriptors. We reconstruct the probability of a correct response as a function of student ability, effectively generating synthetic Item Characteristic Curves (ICCs) to estimate IRT parameters. Evaluation on a dataset of Grade 6 English Language Arts (ELA) items and the BEA 2024 Shared Task dataset demonstrates that this method competes with or outperforms baseline approaches. This simulation-based technique seems particularly effective at modeling item discrimination.
Problem

Research questions and friction points this paper is trying to address.

Item Response Theory
item parameter estimation
field testing
Item Characteristic Curves
psychometric properties
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Item Response Theory
Item Characteristic Curves
Low-Rank Adaptation
Psychometric Modeling
πŸ”Ž Similar Papers
No similar papers found.