Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System

📅 2026-07-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the cross-course scalability of AI teaching assistant systems in real classrooms and examines discrepancies between synthetic agents and actual student behaviors. We introduce LEA, a system integrating course-specific retrieval-augmented generation (RAG) with knowledge component (KC) modeling, supporting three interaction modes: chat, tutoring, and quizzes. LEA was deployed and evaluated across multiple authentic courses for the first time. Results show stable performance in Answer Relevancy (0.88–0.94) and Context Precision (0.88–0.90), yet Faithfulness dropped markedly across courses (0.69→0.50), indicating a bias toward the source course in generated reasoning. The findings demonstrate that synthetic agent evaluations cannot fully predict real-world performance and reveal that certain system components exhibit course dependency, offering critical empirical insights for designing generalizable AI-powered educational systems.
📝 Abstract
This paper is an extension of a paper presented at the ICAART 2026 conference, which introduced LEA (Learning Engagement Assistant), an adaptive AI tutoring agent combining course-specific Retrieval-Augmented Generation (RAG) with structured Knowledge Component (KC) models across integrated Chat, Tutor, and Quiz modes. That prior work validated LEA on a single STEM course (CMP511) exclusively through simulation, using synthetic learner agents. This paper extends that work by reporting the first classroom deployment of LEA with real students (n = 8, CMP511) and the first empirical test of its cross-course scalability, deploying the system across three courses spanning two academic levels and two disciplinary domains. The study reveals a divergence from simulation predictions across modes, showing that synthetic evaluation alone cannot anticipate all aspects of real deployment. A RAGAS-based cross-course scalability evaluation (660 questions) finds Answer Relevancy and Context Precision broadly stable across courses (0.88-0.94 and 0.88-0.90 respectively), while Faithfulness declines with curriculum distance from the system's original course (0.69 to 0.50), a preliminary finding that may reflect generation logic tuned to the system's original subject rather than a scalability limitation. These findings suggest that while the orchestration layer requires no modification, full course-agnosticism of all downstream components requires further investigation.
Problem

Research questions and friction points this paper is trying to address.

AI tutoring system
cross-course scalability
classroom deployment
learning engagement
simulation vs real-world evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Retrieval-Augmented Generation (RAG)
Knowledge Component (KC) models
cross-course scalability
agentic AI tutoring
RAGAS evaluation
T
Teri Rumble
Faculty of Design, Informatics and Business, Abertay University, Dundee, United Kingdom
Javad Zarrin
Javad Zarrin
Senior Lecturer, Abertay University, Dundee, UK
Artificial IntelligenceDistributed SystemsNetworksSchedulingOptimisation
P
P. George Lovell
Faculty of Design, Informatics and Business, Abertay University, Dundee, United Kingdom
R
Ruth Falconer
Faculty of Design, Informatics and Business, Abertay University, Dundee, United Kingdom