🤖 AI Summary
This study addresses the challenge of model selection in model-based reinforcement learning, where distribution shifts between simulators and target environments complicate adaptation, and conventional correction methods often fail under limited target data. To overcome this, we propose the MC-WM framework, which employs a data partitioning strategy to disentangle source and target domain data, introduces normalized calibration risk as a model selection criterion, and dynamically routes imagined updates via confidence-weighted signals. This approach enables adaptation to novel environments without requiring reward function redesign. The effectiveness of MC-WM is validated through 541 evaluations across three distribution shift scenarios in MuJoCo, demonstrating robust environment adaptation under few-shot conditions.
📝 Abstract
Model-based reinforcement learning (MBRL) can exploit simulated experience, but a simulator-to-target shift creates a model-selection problem: correcting the simulator and fitting the target directly can each fail under limited target data. We introduce the Model-Corrected World Model (MC-WM), which separates initial target data into disjoint fit, selection, and calibration partitions and deploys the family with lower standardized calibration risk. A learned confidence signal and deterministic validity predicates weight one-step imagined policy updates without rewriting physical rewards. We evaluate 540 unique reported run cells across three controlled Multi-Joint dynamics with Contact (MuJoCo) shifts; one exact-routing cell was repeated after a pre-deployment artifact gate, giving 541 completed executions.