🤖 AI Summary
This study addresses the challenge of identifying and filtering detrimental trajectories during online reinforcement learning training. To this end, we propose a gradient-based dynamic trajectory valuation framework that eliminates the need for a validation set by efficiently evaluating and filtering low-quality trajectories at the mini-batch level using gradient information. This approach can be seamlessly integrated into mainstream algorithms such as PPO, GRPO, and DPO with minimal computational overhead. Experimental results demonstrate that the proposed framework significantly enhances model performance and data efficiency while effectively improving the stability of the optimization process.
📝 Abstract
We consider the problem of trajectory valuation in reinforcement learning: how to identify and mitigate detrimental trajectories during online training. Unlike classification, where data valuation relies on fixed training and validation sets, reinforcement learning involves dynamically generated trajectories without explicit validation signals, making conventional influence-based methods inapplicable. We propose Dynamic Trajectory Valuation (DTV), a simple and efficient framework that estimates trajectory utility at the mini-batch level and filters detrimental trajectories based solely on gradient information. By operating at the optimization level, DTV integrates seamlessly with existing reinforcement learning pipelines with minimal overhead. Extensive experiments across diverse settings, including PPO, GRPO, and DPO, demonstrate that DTV consistently improves performance, enhances data efficiency, and stabilizes optimization.