From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing in-context reinforcement learning methods rely on behavior cloning, rendering policies susceptible to biases in low-quality offline data. This work proposes the QTPT model, which retains the Transformer architecture while replacing supervised prediction with a Bellman-style Q-objective. By leveraging in-context reward and transition information to estimate action values, QTPT achieves a paradigm shift from imitation to value estimation. Theoretically, we establish its robustness to data quality in stochastic linear bandits and finite-horizon Markov decision processes. Empirically, the proposed method significantly outperforms conventional supervised prediction approaches across reinforcement learning benchmarks involving noisy or suboptimal data, as well as on D4RL Kitchen and AntMaze tasks.
📝 Abstract
Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. This enables task inference from context, but makes the learned policy strongly depend on the quality of offline actions: when trajectories are weak or suboptimal, imitation itself becomes a biased learning signal. We propose Q-Target Pretrained Transformers (QTPT), which keeps the context-conditioned Transformer architecture but replaces behavior cloning with a Bellman-style Q-target objective. QTPT therefore learns to use rewards and transitions in the context to estimate action values, rather than simply imitating the behavior policy. We theoretically analyze QTPT in stochastic linear bandits and finite-horizon MDPs, showing stronger robustness to data quality than supervised pretraining. Empirically, QTPT improves over supervised behavior prediction on controlled RL benchmarks with random or suboptimal data, and we examine extensions to D4RL Kitchen and AntMaze. Supplementary experiments evaluate backbone robustness, meta-RL comparisons, task-coherent context, and unsupported-action value overestimation. These comparisons distinguish the benefits of Q-target pretraining from the remaining limitations of offline coverage.
Problem

Research questions and friction points this paper is trying to address.

In-context reinforcement learning
Suboptimal offline data
Behavior cloning bias
Data quality dependence
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-Context Reinforcement Learning
Q-Target Pretrained Transformers
Bellman Objective
Behavior Cloning
Suboptimal Data
🔎 Similar Papers
Y
Yichen Lin
Shanghai Jiao Tong University, Shanghai, China
X
Xuyuan Xiong
Shanghai Jiao Tong University, Shanghai, China
X
Xue Wang
Alibaba Group
X
Xiangfu Meng
Shanghai Jiao Tong University, Shanghai, China
M
Mike Mingcheng Wei
School of Management, University at Buffalo, Buffalo, NY , USA
Tao Yao
Tao Yao
Alibaba
Operations ResearchMachine LearningAnalyticsStatisticsTransportation