Automating Constructive Assessment with Large Language Models: Toward Scalable and Repeated Evaluation of Practical Competence

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过大型语言模型自动生成和评估包含错误的案例问题,自动评分并提供反馈,以解决HDR方法在评估实践技能时需要专业知识和大量努力的问题。
📝 Abstract
This study aimed to automate hierarchical diagnostic reasoning (HDR), a constructive method for evaluating practical judgment skills, by developing and testing an evaluation process using a large language model. HDR is a descriptive task that measures higher-order cognitive skills by requiring students to identify and explain errors in case-study-based problems. However, it requires expertise and effort to develop and evaluate. Hence, we proposed and empirically validated the automatic (1) generation of case problems containing errors aligned with educational intentions, (2) scoring of descriptive answers, and (3) generation of structured feedback based on incorrect answers, achieved solely through prompt design without fine-tuning. The internal consistency and construct validity of the generated problems were supported by the experimental score distribution and Cronbach's alpha (0.78). The agreement between automated and human ratings reached 100% under some conditions. The feedback was rated as being as convincing and useful as that from human instructors, demonstrating a practical framework for implementing HDR-based constructive assessment with reproducibility, immediacy, and low cost. The flexibility of large language models will also enable repeated and longitudinal assessments while maintaining structure, showing broad potential for application in educational settings.
Problem

Research questions and friction points this paper is trying to address.

Automating Constructive Assessment
Hierarchical Diagnostic Reasoning
Large Language Models
Practical Competence
Innovation

Methods, ideas, or system contributions that make the work stand out.

large language model
hierarchical diagnostic reasoning
automatic feedback generation
constructive assessment
prompt design
🔎 Similar Papers
No similar papers found.
S
Satoshi Takahashi
Graduate School of Economics, Nagoya University, Furo-cho, Chikusa-ku, Nagoya 464-8601, Japan
Atsushi Yoshikawa
Atsushi Yoshikawa
Kanto Gakuin University
M
Megumi Kose
Graduate School of Management, GLOBIS University, Sumitomo Fudosan Kojimachi Building, 5-1 Nibancho, Chiyoda-ku, Tokyo 102-0084, Japan
K
Kenichi Suzuki
GLOBIS AI Management Education Research Institute, Graduate School of Management, GLOBIS University, Sumitomo Fudosan Kojimachi Building, 5-1 Nibancho, Chiyoda-ku, Tokyo 102-0084, Japan
C
Chieko Inoue
Graduate School of Management, GLOBIS University, Sumitomo Fudosan Kojimachi Building, 5-1 Nibancho, Chiyoda-ku, Tokyo 102-0084, Japan
Y
Yumi Watanabe
Graduate School of Management, GLOBIS University, Sumitomo Fudosan Kojimachi Building, 5-1 Nibancho, Chiyoda-ku, Tokyo 102-0084, Japan
M
Mari Sawada
GLOBIS Corporation, Sumitomo Fudosan Kojimachi Building, 5-1 Nibancho, Chiyoda-ku, Tokyo 102-0084, Japan