đ€ AI Summary
This study addresses the limitations of manual grading and rule-based automated scoring in command-line examinationsâspecifically their poor scalability, inability to handle partial credit, equivalent answers, and syntactic variationsâby proposing a large language model (LLM)-based automatic scoring approach. The authors introduce the first four-tier question taxonomy that integrates cognitive complexity and operational impact to delineate the boundaries of AI scoring applicability and design structured prompts to enhance scoring consistency. Experimental results across multiple LLMsâincluding GPT, Claude Opus, Gemini, and GLMâshow that Gemini 1.5 Pro, augmented with rubric-enhanced prompts, achieves the highest humanâAI agreement (ICC = 0.888). The findings demonstrate that question complexity effectively predicts LLM scoring accuracy and establish a transferable evaluation protocol alongside reusable prompt templates.
đ Abstract
Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking difficult and rule-based autograders cannot handle partial credit, equivalent solutions, or syntactic variation. This paper evaluates whether four frontier Large Language Models (GPT, Claude Opus, Gemini, and GLM) can approximate expert judgment when grading short Linux/bash command responses. The study adopts a four-level cognitive taxonomy that combines cognitive complexity and operational impact, ranging from information retrieval (L1) and basic file manipulation (L2) to structural operations (L3) and advanced system management (L4). The models were tested with two prompt variants, a minimal baseline and a rubric-enhanced version, on 1200 real responses from second-year Computer Engineering students independently graded by three expert instructors. Gemini~3.0 Pro with rubric-guided prompting achieved the highest human-AI agreement (ICC(3,1) = 0.888, MAE = 0.10, Bland-Altman bias = -0.014). Agreement declined consistently as taxonomy level increased, with the largest discrepancies at higher levels. Across all models, rubric quality had a larger effect than provider choice, with structured prompts consistently improving agreement. These results show that question complexity is a reliable predictor of the difficulty LLMs face in grading accurately, and they establish a principled, taxonomy-based framework for determining which questions are suitable for AI-assisted grading and which require human review, while also providing a transferable evaluation protocol and prompt templates.