Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors

📅 2026-03-26
📈 Citations: 0
Influential: 0
📄 PDF

career value

162K/year
🤖 AI Summary
This study investigates whether large language model (LLM)-based automated scoring systems are susceptible to construct-irrelevant factors such as spelling errors, textual redundancy, and off-topic content. To this end, a dual-architecture LLM scoring system was developed for automatically evaluating short-answer responses in situational judgment tests, and its robustness was systematically assessed through adversarial perturbation experiments. Results indicate that the system demonstrates strong stability against meaningless padding, spelling errors, and variations in stylistic complexity, while appropriately penalizing repetitive text and off-topic responses. Overall, its scoring performance surpasses that of conventional non-LLM approaches. This work provides the first systematic validation of how LLM-based scoring responds to diverse construct-irrelevant perturbations, revealing its distinctive scoring logic and advantages.

Technology Category

Application Category

📝 Abstract
Automated systems have been widely adopted across the educational testing industry for open-response assessment and essay scoring. These systems commonly achieve performance levels comparable to or superior than trained human raters, but have frequently been demonstrated to be vulnerable to the influence of construct-irrelevant factors (i.e., features of responses that are unrelated to the construct assessed) and adversarial conditions. Given the rising usage of large language models in automated scoring systems, there is a renewed focus on ``hallucinations'' and the robustness of these LLM-based automated scoring approaches to construct-irrelevant factors. This study investigates the effects of construct-irrelevant factors on a dual-architecture LLM-based scoring system designed to score short essay-like open-response items in a situational judgment test. It was found that the scoring system was generally robust to padding responses with meaningless text, spelling errors, and writing sophistication. Duplicating large passages of text resulted in lower scores predicted by the system, on average, contradicting results from previous studies of non-LLM-based scoring systems, while off-topic responses were heavily penalized by the scoring system. These results provide encouraging support for the robustness of future LLM-based scoring systems when designed with construct relevance in mind.
Problem

Research questions and friction points this paper is trying to address.

construct-irrelevant factors
LLM-based scoring
automated essay scoring
robustness
hallucinations
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-based scoring
construct-irrelevant factors
robustness
automated essay scoring
situational judgment test
🔎 Similar Papers
No similar papers found.