In-Distribution Imagination for Model-Based Offline Reinforcement Learning
This study addresses the problem of accumulated errors causing imagined trajectories to deviate from the data distribution in model-based offline reinforcement learning. To this end, it proposes an in-distribution imagination framework that estimates trajectory support within a learned representation space. Innovatively adopting trajectory-level rather than transition-level uncertainty control, the method adaptively truncates rollouts that depart from the data manifold and incorporates entropy regularization to ensure stable model training. Experimental results demonstrate that this framework significantly enhances performance under limited data conditions. Furthermore, the findings confirm that trajectory support serves as a more accurate predictor of rollout failures compared to conventional transition-level uncertainty measures.