Not All Rollouts Are Worth Learning: On Trajectory Valuation for Post-Training Reinforcement Learning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of identifying and filtering detrimental trajectories during online reinforcement learning training. To this end, we propose a gradient-based dynamic trajectory valuation framework that eliminates the need for a validation set by efficiently evaluating and filtering low-quality trajectories at the mini-batch level using gradient information. This approach can be seamlessly integrated into mainstream algorithms such as PPO, GRPO, and DPO with minimal computational overhead. Experimental results demonstrate that the proposed framework significantly enhances model performance and data efficiency while effectively improving the stability of the optimization process.
📝 Abstract
We consider the problem of trajectory valuation in reinforcement learning: how to identify and mitigate detrimental trajectories during online training. Unlike classification, where data valuation relies on fixed training and validation sets, reinforcement learning involves dynamically generated trajectories without explicit validation signals, making conventional influence-based methods inapplicable. We propose Dynamic Trajectory Valuation (DTV), a simple and efficient framework that estimates trajectory utility at the mini-batch level and filters detrimental trajectories based solely on gradient information. By operating at the optimization level, DTV integrates seamlessly with existing reinforcement learning pipelines with minimal overhead. Extensive experiments across diverse settings, including PPO, GRPO, and DPO, demonstrate that DTV consistently improves performance, enhances data efficiency, and stabilizes optimization.
Problem

Research questions and friction points this paper is trying to address.

Trajectory Valuation
Reinforcement Learning
Post-Training
Data Valuation
Detrimental Trajectories
Innovation

Methods, ideas, or system contributions that make the work stand out.

Trajectory Valuation
Reinforcement Learning
Gradient-based Filtering
Post-Training
Data Efficiency
X
Xuesong Jia
Department of Computer Science, Brandeis University
Z
Ziao Yang
Department of Computer Science, Brandeis University
Z
Zhanhe Huang
Department of Computer Science, Brandeis University
Hongfu Liu
Hongfu Liu
Brandeis University