Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach

📅 2026-07-02
📈 Citations: 0
✹ Influential: 0
📄 PDF
đŸ€– AI Summary
This study addresses the limitations of manual grading and rule-based automated scoring in command-line examinations—specifically their poor scalability, inability to handle partial credit, equivalent answers, and syntactic variations—by proposing a large language model (LLM)-based automatic scoring approach. The authors introduce the first four-tier question taxonomy that integrates cognitive complexity and operational impact to delineate the boundaries of AI scoring applicability and design structured prompts to enhance scoring consistency. Experimental results across multiple LLMs—including GPT, Claude Opus, Gemini, and GLM—show that Gemini 1.5 Pro, augmented with rubric-enhanced prompts, achieves the highest human–AI agreement (ICC = 0.888). The findings demonstrate that question complexity effectively predicts LLM scoring accuracy and establish a transferable evaluation protocol alongside reusable prompt templates.
📝 Abstract
Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking difficult and rule-based autograders cannot handle partial credit, equivalent solutions, or syntactic variation. This paper evaluates whether four frontier Large Language Models (GPT, Claude Opus, Gemini, and GLM) can approximate expert judgment when grading short Linux/bash command responses. The study adopts a four-level cognitive taxonomy that combines cognitive complexity and operational impact, ranging from information retrieval (L1) and basic file manipulation (L2) to structural operations (L3) and advanced system management (L4). The models were tested with two prompt variants, a minimal baseline and a rubric-enhanced version, on 1200 real responses from second-year Computer Engineering students independently graded by three expert instructors. Gemini~3.0 Pro with rubric-guided prompting achieved the highest human-AI agreement (ICC(3,1) = 0.888, MAE = 0.10, Bland-Altman bias = -0.014). Agreement declined consistently as taxonomy level increased, with the largest discrepancies at higher levels. Across all models, rubric quality had a larger effect than provider choice, with structured prompts consistently improving agreement. These results show that question complexity is a reliable predictor of the difficulty LLMs face in grading accurately, and they establish a principled, taxonomy-based framework for determining which questions are suitable for AI-assisted grading and which require human review, while also providing a transferable evaluation protocol and prompt templates.
Problem

Research questions and friction points this paper is trying to address.

automated grading
Linux/bash examinations
large language models
cognitive taxonomy
command-line assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

cognitive taxonomy
LLM-based grading
rubric-enhanced prompting
Linux/bash assessment
human-AI agreement
🔎 Similar Papers
No similar papers found.
đŸ’Œ Related Jobs
No related jobs found.
M
Manuel Alonso-Carracedo
aUniversidade de Vigo, Department of Computer Science, ESEI-Higher School of Computer Engineering, Edificio Politécnico, Campus Universitario As Lagoas s/n, Ourense, 32004, Spain; bIFCAE-Institute for Research in Physics, Computing and Aerospace Science, Universidade de Vigo, Ourense, 32004, Spain
R
Ruben Fernandez-Boullon
aUniversidade de Vigo, Department of Computer Science, ESEI-Higher School of Computer Engineering, Edificio Politécnico, Campus Universitario As Lagoas s/n, Ourense, 32004, Spain; bIFCAE-Institute for Research in Physics, Computing and Aerospace Science, Universidade de Vigo, Ourense, 32004, Spain
P
Pedro Celard
aUniversidade de Vigo, Department of Computer Science, ESEI-Higher School of Computer Engineering, Edificio Politécnico, Campus Universitario As Lagoas s/n, Ourense, 32004, Spain; bIFCAE-Institute for Research in Physics, Computing and Aerospace Science, Universidade de Vigo, Ourense, 32004, Spain
F
Francisco J. Rodriguez-Martinez
aUniversidade de Vigo, Department of Computer Science, ESEI-Higher School of Computer Engineering, Edificio Politécnico, Campus Universitario As Lagoas s/n, Ourense, 32004, Spain; bIFCAE-Institute for Research in Physics, Computing and Aerospace Science, Universidade de Vigo, Ourense, 32004, Spain
Lorena Otero-Cerdeira
Lorena Otero-Cerdeira
University of Vigo
Intelligent AgentsOntology Matching