EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

πŸ“… 2026-07-20
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing methods for automated assessment of instructional videos suffer from limited scalability, inadequate integration of multimodal evidence, and insufficient consideration of learner variability. This work proposes a rubric-based, learner-centered tri-agent large language model evaluation framework that employs a decomposed multi-agent architecture to collaboratively analyze instructional content across multiple dimensions, yielding interpretable, reliable, and complementary assessments. By integrating rubric-guided reasoning, learner role modeling, and multimodal analysis, the approach achieves evaluation reliability on par with the median human expert. Furthermore, its generated feedback reduces expert scoring error from 0.87 to 0.73 and enables experts to effectively identify unreliable model outputs (AUC = 0.77), thereby balancing expert-level accuracy with effective human–AI collaboration.
πŸ“ Abstract
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be evaluated with respect to the intended learner rather than as a universal property. We present EduPanel, a rubric-grounded, learner-conditioned LLM judge that decomposes evaluation across specialized agents to produce interpretable assessments for different aspects of teaching quality. Across expert studies, architecture ablations, and learner-persona analyses, EduPanel achieves reliability comparable to a median human expert. In expert evaluation, its feedback improves scoring accuracy (MAE 0.87 to 0.73), while experts remain able to detect unreliable outputs (AUC = 0.77) instead of accepting them blindly. These results suggest that EduPanel can serve as effective assistants for educational evaluation rather than replacements for human experts.
Problem

Research questions and friction points this paper is trying to address.

teaching video evaluation
pedagogical quality
learner-conditioned assessment
multimodal evidence
scalable evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-agent LLM
learner-conditioned evaluation
interpretable assessment
human-AI collaboration
teaching video quality