In-Distribution Imagination for Model-Based Offline Reinforcement Learning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of accumulated errors causing imagined trajectories to deviate from the data distribution in model-based offline reinforcement learning. To this end, it proposes an in-distribution imagination framework that estimates trajectory support within a learned representation space. Innovatively adopting trajectory-level rather than transition-level uncertainty control, the method adaptively truncates rollouts that depart from the data manifold and incorporates entropy regularization to ensure stable model training. Experimental results demonstrate that this framework significantly enhances performance under limited data conditions. Furthermore, the findings confirm that trajectory support serves as a more accurate predictor of rollout failures compared to conventional transition-level uncertainty measures.
📝 Abstract
Model-based offline reinforcement learning (MBORL) improves sample efficiency through model-generated trajectories. However, accumulative model error can drive imagined trajectories outside the offline data distribution, leading to unrealistic synthetic data and unstable policy optimization. Many existing methods primarily control rollouts using transition-level uncertainty. We propose \emph{in-distribution imagination} (IDI), a rollout control framework that estimates trajectory support in a learned representation space and adaptively truncates rollouts that leave the offline trajectory manifold. Combined with trajectory-regularized RL, an extension of entropy-regularized RL, IDI consistently improves performance in limited-data settings. Experiments show that trajectory support predicts rollout failure substantially better than transition-level uncertainty, highlighting the importance of trajectory-level rollout control in MBORL.
Problem

Research questions and friction points this paper is trying to address.

Model-Based Offline Reinforcement Learning
In-Distribution Imagination
Rollout Control
Trajectory Support
Model Error
Innovation

Methods, ideas, or system contributions that make the work stand out.

Model-Based Offline Reinforcement Learning
In-Distribution Imagination
Trajectory Support
Rollout Control
Trajectory-Regularized RL
🔎 Similar Papers
2024-05-23Trans. Mach. Learn. Res.Citations: 0