The Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning

πŸ“… 2026-10-05
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the cognitive offloading problem in large language model (LLM) tutoring, where reward mechanisms inadvertently encourage directly revealing answers. To mitigate this, we propose a multi-turn reinforcement learning optimization framework. Methodologically, we introduce masked near-transfer post-tests to evaluate genuine instructional efficacy and design a binary reward gating mechanism that replaces continuous penalties to suppress cognitive offloading. The model is further fine-tuned by integrating simulated student interactions with near-transfer datasets. The resulting open-source model, Eduardo-27B, achieves performance comparable to frontier models on standard benchmarks while significantly reducing inference token consumption. This work provides an efficient paradigm for optimizing LLM tutor strategies.
πŸ“ Abstract
Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. The reward is then easiest to raise by telling the student the answer, and a tuned penalty is needed to reduce telling. Drawing on learning sciences, we introduce a masked near-transfer post-test: the student is tested on an unseen variant of the tutored problem with the tutor's utterances masked, so the reward can rise only through what the student wrote in its own turns. This discourages cognitive offloading by the student and allows the continuous penalty to be replaced by two binary reward gates (factual correctness of tutor response, no solution handover). A leave-one-out ablation shows that the learning-gain reward on its own does not separate teaching from telling: the gates reduce solution handover while the near-transfer post-test improves out-of-domain transfer. Using these reward designs we develop Eduardo, a multi-turn RL recipe for training LLM tutors, and use it to train 4B, 9B, 14B and 27B models from two distinct LLM architectures. Our post-trained Eduardo-27B model matches Gemini-3.1-Pro on MathTutorBench and Claude Opus 4.8 on TutorMoments at 2.4-6.2x fewer thinking tokens than frontier models, which matters for interactive tutoring. Without being named in the reward, the model more than doubles its use of the push-for-justification teacher move while support fading (e.g., assigning independent work), whose payoff lies beyond a single-problem dialog episode, is trained out. We open-source our training environment, an 8,671-problem near-transfer dataset, and trained models for further development.
Problem

Research questions and friction points this paper is trying to address.

Assistance Dilemma
Large Language Models
Reinforcement Learning
AI Tutoring
Cognitive Offloading
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-turn Reinforcement Learning
Masked Near-transfer Post-test
Reward Gating
LLM Tutoring
Cognitive Offloading
πŸ”Ž Similar Papers
No similar papers found.