Score
Designs and analyzes algorithms and controllers that use an external teacher or safety module to intermittently override or guide a learner’s actions, producing mixed teacher–student policies and behavior samples. Work in this skill specifies when and how strongly to intervene (including progressive fading), methods to restrain risky exploration via interventions, and techniques to characterize or bound returns and safety for the resulting mixed policies.
This study systematically investigates the feedback-to-update mechanism in on-policy distillation (OPD), where data are generated by the current policy. Framing OPD as a feedback-to-update problem, the work introduces a formula-driven categorization framework that unifies two major update pathways: distributional loss and policy gradient–style log-ratio updates. It further incorporates novel perspectives from temporal credit assignment and temporal vocabulary routing. By leveraging techniques such as KL divergence orientation, generalized advantage estimation (GAE), and counterfactual routing, the analysis reveals that OPD performance critically depends on state compatibility and support set construction. The paper establishes a comprehensive analytical framework for OPD, derives explicit bias bounds, and proposes new methods—GAE-OPD and CR-OPD—to enhance training stability, alongside actionable diagnostic tools and a practical implementation checklist.
This study addresses the absence of a formal definition of instructional safety in existing reinforcement learning–based intelligent tutoring systems, which renders them susceptible to reward hacking—where agents optimize superficial metrics at the expense of genuine learning outcomes. To remedy this, the work introduces a novel four-layer instructional safety model encompassing structural, progression, behavioral, and alignment safety, along with a Reward Hacking Severity Index (RHSI) to quantify goal misalignment. Empirical results from 18,000 learner-agent interactions demonstrate that constraint-based architectures—such as prerequisite enforcement and minimum cognitive demand requirements—reduce RHSI from 0.317 to 0.102, substantially curbing low-value repetitive behaviors, with behavioral safety mechanisms proving most critical. Moreover, constraint-based approaches outperform multi-objective reward designs in ensuring instructional alignment.
This study addresses the challenge in online policy distillation where teacher intervention alters the student's distribution, making it difficult for fixed-intensity strategies to balance trajectory quality against off-policy deviation. To overcome this limitation, this work proposes MAESTRO, a framework that adaptively regulates both the timing of teacher takeover and the generation duration through a segment-aggregated policy divergence scoring mechanism, thereby achieving joint dynamic adaptation of intervention depth and length. Empirical evaluations demonstrate that MAESTRO attains state-of-the-art average accuracy across eight mathematical reasoning benchmarks. Furthermore, it reduces training response lengths by 67.3% compared to standard methods, substantially improving knowledge transfer efficiency.
This study addresses the problem of excessive intervention by LLM tutors, which arises from conflating tutoring capability with intervention necessity. To mitigate this, we propose the DICE framework, which decouples decision-making from generation through explicit selection of pedagogical actions. Furthermore, it introduces a counterfactual metric, Intervention Value (IV), to quantify the utility of interventions and guide policy learning. By integrating IV weighting with KL regularization, the framework optimizes a selective intervention strategy, supported by a multi-variant benchmark, DICE-Bench, for conversation-level evaluation. Experimental results demonstrate that DICE reduces the over-intervention rate to near zero, guiding students toward correct solutions in three to four fewer dialogue turns on average compared to baseline methods.
This study addresses the dual challenges of the “safety–guidance gap” and the “scaffolding paradox” faced by social robots in simultaneously alleviating interview anxiety and enhancing interview skills. Through a three-phase iterative design, the authors propose an Adaptive Scaffolding Ecosystem framework that integrates person-centered therapy, cognitive load theory, and user feedback analysis to establish a dynamic interaction mechanism centered on user agency. This approach enables real-time balancing between emotional support and instructional challenge. Empirical evaluation demonstrates that the system significantly improves users’ psychological safety, engagement, and learning outcomes during mock interviews, offering a novel paradigm for adaptive design in intelligent tutoring robots.
This study addresses the cold-start problem caused by sparse rewards in reinforcement learning for long-horizon agents by proposing the GATS method. GATS employs online policy distillation to provide token-level guidance signals and introduces a novel finding that distillation benefits depend on the performance gap between teacher and student models. Based on this insight, an adaptive weight scheduling mechanism is designed to automatically attenuate and eventually remove teacher guidance as the gap narrows, enabling a smooth transition from guided learning to autonomous exploration while supporting teacher models smaller than the student. Evaluated on benchmarks such as ALFWorld, GATS improves success rates by 4.37%–11.87% over GRPO and achieves the best average performance across all configurations.
This study addresses the challenge in multi-turn agent online distillation where early student errors accumulate, causing trajectory deviation from the teacher distribution and rendering supervision ineffective. To overcome this, we propose the STI-OPD framework, which introduces a novel stochastic intervention mechanism based on policy divergence to replace fixed thresholds, dynamically substituting student actions to balance control with exploration. Furthermore, an importance-weighted reverse KL loss is incorporated to correct sampling bias in mixed trajectories, thereby enhancing optimization reliability. Comprehensive evaluations on tool-use reasoning and long-horizon interaction benchmarks demonstrate that our approach consistently outperforms state-of-the-art baselines, validating the effectiveness of both the stochastic intervention strategy and the importance weighting module.
This study addresses the misalignment between supervision signals and final rewards caused by local distribution shifts in multi-turn agent policy distillation. To overcome this limitation, we propose Outcome-Guided Online Policy Distillation (OG-OPD), which transcends conventional local optimization by introducing a dynamic supervision mechanism based on relative trajectory weights. Specifically, OG-OPD calibrates teacher guidance signals using ultimate task outcomes, enabling precise intervention at critical steps essential for improving success rates. Experimental results demonstrate that OG-OPD significantly outperforms existing baselines on benchmarks such as ALFWorld, achieving an improvement in task success rate of up to 17.7 percentage points.
This study addresses the high computational cost of teacher supervision in online knowledge distillation by proposing an adaptive supervision selection mechanism that leverages the student model's own successful trajectories as a reference. Specifically, this method precisely allocates teacher resources by filtering out failed samples, while integrating hidden state trajectory discrepancy analysis with a budget-aware token selection strategy to achieve efficient supervision. Experimental results demonstrate that the proposed approach maintains comparable inference performance while requiring only 3.46% to 5.02% of the teacher inputs, thereby substantially reducing the computational overhead associated with online distillation.
This study addresses the challenges of sparse rewards limiting exploration in agent reinforcement learning and the difficulty of fixed distillation strategies in dynamically balancing teacher guidance with reward optimization. To this end, we propose TIDE, a method that innovatively leverages the divergence trend between teacher and student policies as a scheduling signal. Through global adaptive scheduling, TIDE enables a smooth transition from distillation to reinforcement learning, while locally fine-tuning multi-turn interaction update weights based on action values and divergence degrees. By integrating the GRPO algorithm with online policy distillation, experimental results demonstrate that TIDE significantly outperforms conventional fixed mixing strategies across multiple benchmarks and models of varying scales, effectively enhancing agents' long-horizon interaction capabilities and overall performance.
This study addresses the limitation of existing agent training methods that rely on fixed scenarios and lack dynamic curriculum adjustment mechanisms aligned with evolving capabilities. To this end, we propose an automated curriculum learning framework that enables the co-evolution of training scenarios and agents. This framework introduces curriculum adaptability as a core dimension for the first time, employing a non-stationary multi-armed bandit model to abstract reusable failure modes, quantify learning progress, and dynamically update training priorities. Experimental results demonstrate that our method improves Pass@1 by 4.4 to 7.5 percentage points on benchmarks such as GAIA2, significantly enhancing overall agent performance.