EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of existing evaluations for educational large language models (LLMs), which typically focus on single-turn interactions and fail to assess sustained pedagogical relationships. The authors propose the first 30-day benchmark designed for continuous teaching, leveraging real student data to train a knowledge tracing model that drives simulated learners. The framework evaluates teacher agents across 55 scenarios on learning gains, responsiveness, and helpfulness, incorporating instructional design principles into a multidimensional scoring system. Validation through multiple LLM-based judges, low expected calibration error (ECE = 0.049), and in-classroom studies demonstrates high behavioral fidelity between simulated and real learners. Results reveal that teaching effectiveness is jointly determined by the base LLM and the agent adapter, with most combinations failing to sustain high-quality tutoring throughout the full duration—providing a foundation for developing trustworthy AI teaching agents.
📝 Abstract
Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point solutions been integrated into agents operating over a learning management system (LMS). Yet tutoring is long-horizon, since a learner improves over days and weeks rather than in a single turn, and no benchmark evaluates an agent tutor across a sustained relationship. We introduce EduClaw-Bench, a benchmark that places an agent tutor in a continuous 30-day relationship with a simulated learner grounded in knowledge tracing (KT), whose knowledge-concept mastery, from a KT model trained on real-student data, drives its answers and is probed for learning gain across 55 scenarios. Each agent is scored on three primary axes (learning gain, responsiveness, and helpfulness) and two curriculum-design axes (Gagné and Rosenshine), with helpfulness and the curriculum axes judged by a cross-family panel of three LLM judges. Evaluating 10 agent adapters over three base-model tiers yields two findings that single-tier, single-session evaluation cannot reach. First, tutoring quality belongs to the base model and the agent harness together rather than either alone. Second, almost no combination sustains good tutoring over the full horizon. A calibration check ($\text{ECE}=0.049$) and a live-classroom field study confirm that the simulated learner and its measurements track reality. Our work is a step toward trustworthy AI tutors for future education.
Problem

Research questions and friction points this paper is trying to address.

long-horizon tutoring
pedagogical agents
learning gain
simulated learners
educational benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

long-horizon benchmark
pedagogical LLM agents
simulated learners
knowledge tracing
AI tutoring evaluation
🔎 Similar Papers
No similar papers found.