🤖 AI Summary
This study addresses the problem of excessive intervention by LLM tutors, which arises from conflating tutoring capability with intervention necessity. To mitigate this, we propose the DICE framework, which decouples decision-making from generation through explicit selection of pedagogical actions. Furthermore, it introduces a counterfactual metric, Intervention Value (IV), to quantify the utility of interventions and guide policy learning. By integrating IV weighting with KL regularization, the framework optimizes a selective intervention strategy, supported by a multi-variant benchmark, DICE-Bench, for conversation-level evaluation. Experimental results demonstrate that DICE reduces the over-intervention rate to near zero, guiding students toward correct solutions in three to four fewer dialogue turns on average compared to baseline methods.
📝 Abstract
Fluent guidance is not the same as useful intervention. LLM tutors are typically trained to generate the next teacher utterance, implicitly assuming that every student turn warrants a response. However, our experiments indicate that this conflates tutoring capability (what to say) with intervention necessity (whether to say it). We introduce DICE, a framework that decouples intervention decisions from response generation by first selecting an explicit pedagogical action. To calibrate this action selection policy, we define Intervention Value (IV), a rollout-grounded counterfactual metric that compares each action against non-intervention. IV shows that many prescribed interventions provide little or no marginal benefit. We further introduce DICE-Bench, a multi-variant math tutoring benchmark with skill-preserving problem variants for session-level evaluation. Using IV-weighted and KL-regularized policy optimization, DICE learns to intervene selectively while preserving tutoring effectiveness. In simulated tutoring sessions, DICE reduces the over-intervention rate to near zero while guiding students to correct solutions in approximately 3-4 fewer turns on average than existing Socratic tutoring baselines.