Q-learning Penalized Transformer for Safe Offline Reinforcement Learning

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of harmonizing constraint satisfaction, reward optimization, and behavior regularization in safe offline reinforcement learning by proposing the QPT framework. This approach integrates conditional sequence modeling with constraint-aware value estimation, introducing a Q-type penalty mechanism that injects explicit safety semantics into the Transformer policy generation process. Consequently, it achieves consistent safe policy optimization across training and inference, effectively closing the loop between the learning and deployment phases. Experimental results demonstrate that QPT comprehensively outperforms existing baseline methods across 38 tasks on the DSRL benchmark, while exhibiting robust zero-shot generalization capabilities under varying constraint thresholds.
📝 Abstract
This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing three highly interconnected and competing objectives: satisfying safety constraints, maximizing rewards, and adhering to the behavior regularization imposed by the offline dataset. To tackle this trilogy challenge, we propose Q-learning Penalized Transformer policy (QPT), a \emph{training--inference consistent} framework that bridges conditional sequence modeling with constraint-aware value estimation. QPT trains a Transformer policy that generates actions conditioned on trajectory context and target return/cost, retaining strong behavior regularization. To inject explicit safety semantics during learning, we augment sequence-model training with a Q-shaped penalty using learned reward and cost Q-functions to favor high return under low constraint violation. At inference, the same Q-functions enforce the cost threshold and choose the highest-reward feasible action, closing the loop between training and deployment. We provide a principled analysis under stylized near-deterministic CMDPs, characterizing how Q-penalized conditional generation improve safety and performance. Empirically, QPT consistently outperforms strong safe offline RL baselines across 38 tasks on the DSRL benchmark, and exhibits robust zero-shot adaptation to different constraint thresholds.
Problem

Research questions and friction points this paper is trying to address.

safe offline reinforcement learning
safety constraints
behavior regularization
constrained Markov decision processes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Safe Offline Reinforcement Learning
Q-learning Penalized Transformer
Conditional Sequence Modeling
Training-Inference Consistency
Zero-shot Adaptation