VIGOR: Zero-Shot Visual Generalization via Latent-Space Consistency in Model-Based Reinforcement Learning

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the severe performance degradation of model-based reinforcement learning under unseen visual distractions caused by error accumulation. To this end, we propose VIGOR, a framework that effectively blocks the propagation of visual perturbations during recurrent planning through asymmetric data augmentation, a novel latent dynamics consistency mechanism, and encoder drift suppression. VIGOR achieves zero-shot visual generalization while maintaining high sample efficiency, and demonstrates robustness to specific augmentation strategies. Experimental evaluations on the DeepMind Control Suite and Robosuite benchmarks show that VIGOR outperforms existing state-of-the-art methods by 3.4% and 43.6%, respectively, validating its superior effectiveness and robustness in visually complex environments.
📝 Abstract
Model-based reinforcement learning (MBRL) achieves strong sample efficiency by planning within learned latent dynamics, yet its performance degrades substantially under unseen visual distractions such as background variations, lighting changes, or camera shifts. Unlike model-free RL, where encoder perturbations affect only single-step predictions, MBRL suffers from a two-level vulnerability: visual distractions first push encoder outputs out of distribution, and these errors then compound through recursive latent rollouts over the planning horizon. We propose visual generalization via latent-space consistency in model-based RL (VIGOR), a framework that enables zero-shot generalization to unseen visual distractions while retaining the sample efficiency of its MBRL backbone. VIGOR integrates three interdependent components: (i) asymmetric weak-to-strong augmentation, which pairs weak-only and weak-to-strong latent views within a single batch; (ii) dynamics-level consistency, which enforces augmentation-invariant transition predictions through direct latent regression; and (iii) encoder-level stabilization, which prevents encoder drift under the cross-augmentation supervision imposed by dynamics-level consistency. Evaluations on the DeepMind Control Suite (DMC) and Robosuite show that VIGOR outperforms state-of-the-art model-free and model-based baselines, surpassing the second-best baseline by 3.4% on DMC and 43.6% on Robosuite. Ablations further show that VIGOR's robustness is augmentation-agnostic: replacing the default augmentation with alternatives from distinct perturbation families preserves strong generalization, confirming that latent-space consistency, not the augmentation choice, drives robustness.
Problem

Research questions and friction points this paper is trying to address.

Model-Based Reinforcement Learning
Visual Generalization
Visual Distractions
Zero-Shot Generalization
Latent Dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Model-Based Reinforcement Learning
Zero-Shot Visual Generalization
Latent-Space Consistency
Data Augmentation
Encoder Stabilization
🔎 Similar Papers