A Bandit-Based Approach to Educational Recommender Systems: Contextual Thompson Sampling for Learner Skill Gain Optimization

πŸ“… 2026-02-04
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of dynamically adapting personalized exercise recommendations to learners’ evolving skill levels in large-scale online education. The authors propose a contextual Thompson sampling-based multi-armed bandit approach that integrates learner characteristics and historical performance to perform real-time Bayesian inference of skill proficiency. By optimizing exercise sequences with the explicit objective of maximizing skill gain, the method tailors recommendations to individual learning trajectories. Experiments on real-world data from a mathematics tutoring platform demonstrate that the proposed approach significantly enhances skill acquisition, effectively accommodates individual differences, and simultaneously identifies high-impact exercises and at-risk learners requiring intervention. The solution exhibits strong scalability and practical utility for real-world educational applications.

Technology Category

Machine Learning: Online Learning & BanditsSearch and Optimization: Learning to SearchReasoning under Uncertainty: Stochastic Optimization

Application Category

User Modeling, Personalization and Recommendation: Practical large-scale studies of user experienceSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: Economics and fairness of platforms and recommendation systems
πŸ“ Abstract
In recent years, instructional practices in Operations Research (OR), Management Science (MS), and Analytics have increasingly shifted toward digital environments, where large and diverse groups of learners make it difficult to provide practice that adapts to individual needs. This paper introduces a method that generates personalized sequences of exercises by selecting, at each step, the exercise most likely to advance a learner's understanding of a targeted skill. The method uses information about the learner and their past performance to guide these choices, and learning progress is measured as the change in estimated skill level before and after each exercise. Using data from an online mathematics tutoring platform, we find that the approach recommends exercises associated with greater skill improvement and adapts effectively to differences across learners. From an instructional perspective, the framework enables personalized practice at scale, highlights exercises with consistently strong learning value, and helps instructors identify learners who may benefit from additional support.
Problem

Research questions and friction points this paper is trying to address.

educational recommender systems
personalized learning
skill gain optimization
adaptive practice
learner modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contextual Thompson Sampling
Educational Recommender Systems
Personalized Learning
Skill Gain Optimization
Bandit Algorithms
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
L
Lukas De Kerpel
Faculty of Economics and Business Administration, Ghent University; FlandersMake@UGent, corelab CV AMO, Tweekerkenstraat 2, 9000, Ghent, Belgium
A
Arthur Thuy
Faculty of Economics and Business Administration, Ghent University; FlandersMake@UGent, corelab CV AMO, Tweekerkenstraat 2, 9000, Ghent, Belgium
Dries F. Benoit
Dries F. Benoit
Associate professor of Data Analytics, Ghent University
Data ScienceMachine LearningBayesian Statistics