🤖 AI Summary
Existing in-context reinforcement learning methods rely on behavior cloning, rendering policies susceptible to biases in low-quality offline data. This work proposes the QTPT model, which retains the Transformer architecture while replacing supervised prediction with a Bellman-style Q-objective. By leveraging in-context reward and transition information to estimate action values, QTPT achieves a paradigm shift from imitation to value estimation. Theoretically, we establish its robustness to data quality in stochastic linear bandits and finite-horizon Markov decision processes. Empirically, the proposed method significantly outperforms conventional supervised prediction approaches across reinforcement learning benchmarks involving noisy or suboptimal data, as well as on D4RL Kitchen and AntMaze tasks.
📝 Abstract
Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. This enables task inference from context, but makes the learned policy strongly depend on the quality of offline actions: when trajectories are weak or suboptimal, imitation itself becomes a biased learning signal. We propose Q-Target Pretrained Transformers (QTPT), which keeps the context-conditioned Transformer architecture but replaces behavior cloning with a Bellman-style Q-target objective. QTPT therefore learns to use rewards and transitions in the context to estimate action values, rather than simply imitating the behavior policy. We theoretically analyze QTPT in stochastic linear bandits and finite-horizon MDPs, showing stronger robustness to data quality than supervised pretraining. Empirically, QTPT improves over supervised behavior prediction on controlled RL benchmarks with random or suboptimal data, and we examine extensions to D4RL Kitchen and AntMaze. Supplementary experiments evaluate backbone robustness, meta-RL comparisons, task-coherent context, and unsupported-action value overestimation. These comparisons distinguish the benefits of Q-target pretraining from the remaining limitations of offline coverage.