🤖 AI Summary
This study addresses the policy gradient bias induced by value function approximation errors in Actor-Critic methods. We propose an offline-to-online reinforcement learning framework based on advantage variance reduction. By revealing the intrinsic connection between covariance structure and reward decomposition, our approach integrates control variates with direct advantage estimation to construct an Actor-Critic architecture that combines offline pretraining with online fine-tuning, thereby achieving efficient and unbiased policy optimization. Experimental results demonstrate that the proposed method attains performance comparable to GRPO on mathematical reasoning tasks while substantially reducing the number of online interaction steps, leading to significantly lower computational costs.
📝 Abstract
Recent progress in reinforcement learning with verifiable rewards (RLVR) has highlighted the effectiveness of simple critic-free policy-gradient methods such as Group Relative Policy Optimization (GRPO). In contrast, actor-critic methods rely on learned value functions whose approximation error can introduce bias through commonly used advantage estimators such as temporal-difference error. Motivated by this observation, we revisit trajectory-level control variates through an advantage-value formulation, which we call Advantage-Based Control Variates (ABC). This formulation reveals that the covariance structure is closely related to the return decomposition used in Direct Advantage Estimation (DAE). Finally, we combine ABC with DAE into a single actor-critic algorithm and evaluate it in an offline-to-online RLVR setting, where the critic is first trained on previously collected trajectories and adapted during online learning. On mathematical reasoning tasks, ABC achieves performance competitive with GRPO using substantially fewer online optimization steps.