Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

πŸ“… 2026-07-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the unclear efficacy of offline pretraining for the Q-function in online reinforcement learning fine-tuning under pretrained policy strategies. It reveals a fundamental mismatch between the objectives of Q-function pretraining and online fine-tuning, which limits performance gains. To overcome this issue, the paper proposes Initialization via Policy Ensemble (IPE), a method that leverages rollout data from an ensemble of diverse policies to guide the initialization of the Q-function, thereby enabling more effective knowledge transfer. Evaluated across multiple continuous control benchmarks, IPE achieves an average 1.26Γ— improvement in fine-tuning performance over naive Q-function pretraining, substantially enhancing online learning efficiency.
πŸ“ Abstract
Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too? Conventional wisdom suggests it should, but recent results show that online RL with a randomly-initialized Q-function can result in highly performant and reliable policies without needing to pretrain the Q-function. In this paper, we systematically study whether pretraining the Q-function actually helps when fine-tuning on top of a pretrained base policy. We find, surprisingly, that naive Q-function pretraining often provides little benefit over random initialization. We show this stems from a fundamental mismatch: the Q-function learned during pretraining targets the pretrained policy's Q-function, not the Q-function that online fine-tuning converges to, and this gap persists even after offline value maximization. Motivated by this finding, we propose Initialization via Policy Ensemble (IPE), a simple method that trains multiple diverse policies and uses their pooled rollouts to bootstrap the Q-function learning in online RL. Across a suite of challenging continuous control benchmarks, IPE yields an average 1.26x improvement in fine-tuning performance over naive Q-function pre-training.
Problem

Research questions and friction points this paper is trying to address.

Q-function pretraining
online RL fine-tuning
value-based reinforcement learning
policy pretraining
offline-to-online transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

Q-function pretraining
online reinforcement learning
policy ensemble
fine-tuning
value-based RL
πŸ”Ž Similar Papers
No similar papers found.