Unsupervised Data Generation for Offline Reinforcement Learning: A Perspective from Model

📅 2025-06-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Offline reinforcement learning (RL) suffers from poor policy generalization due to distributional shift in batch data, especially under task-agnostic settings where performance degrades significantly. To address this, we establish a theoretical link between batch data quality and algorithmic performance from a model-centric perspective—proving for the first time that the distributional distance between the behavior policy and the optimal policy is a fundamental determinant of performance gap. Building on this insight, we propose Unsupervised Data Generation (UDG), a novel framework that automatically optimizes data distributions by minimizing worst-case regret, without requiring task labels or expert demonstrations. UDG integrates model-based offline RL, state-action distribution modeling, and regret minimization theory, featuring dual mechanisms for data generation and selection. Empirical evaluation across multiple benchmarks demonstrates that UDG consistently outperforms supervised data generation methods, substantially improving both policy performance and robustness on unseen tasks.

Technology Category

Machine Learning: Online Learning & BanditsGame Theory and Economic Paradigms: Adversarial LearningSearch and Optimization: Learning to Search

Application Category

Economics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Offline reinforcement learning (RL) recently gains growing interests from RL researchers. However, the performance of offline RL suffers from the out-of-distribution problem, which can be corrected by feedback in online RL. Previous offline RL research focuses on restricting the offline algorithm in in-distribution even in-sample action sampling. In contrast, fewer work pays attention to the influence of the batch data. In this paper, we first build a bridge over the batch data and the performance of offline RL algorithms theoretically, from the perspective of model-based offline RL optimization. We draw a conclusion that, with mild assumptions, the distance between the state-action pair distribution generated by the behavioural policy and the distribution generated by the optimal policy, accounts for the performance gap between the policy learned by model-based offline RL and the optimal policy. Secondly, we reveal that in task-agnostic settings, a series of policies trained by unsupervised RL can minimize the worst-case regret in the performance gap. Inspired by the theoretical conclusions, UDG (Unsupervised Data Generation) is devised to generate data and select proper data for offline training under tasks-agnostic settings. Empirical results demonstrate that UDG can outperform supervised data generation on solving unknown tasks.
Problem

Research questions and friction points this paper is trying to address.

Addresses out-of-distribution issue in offline RL
Links batch data quality to offline RL performance
Proposes unsupervised data generation for task-agnostic settings
Innovation

Methods, ideas, or system contributions that make the work stand out.

Model-based offline RL optimization perspective
Unsupervised RL minimizes worst-case regret
UDG generates task-agnostic offline training data
💼 Related Jobs
No related jobs found.
S
Shuncheng He
Tsinghua University
H
Hongchang Zhang
Tsinghua University
Jianzhun Shao
Jianzhun Shao
Alibaba Inc.
reinforcement learning in LLMmulti-agent reinforcement learningoffline reinforcement learning
Y
Yuhang Jiang
Tsinghua University
X
Xiangyang Ji
Tsinghua University