FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the low sample efficiency of visual reinforcement learning and the significant modality gap between privileged simulation states and real-world RGB-D inputs. To overcome these challenges, this work proposes a single-stage Proximal Policy Optimization (PPO) framework featuring a novel automatic modality regulation mechanism based on action distribution discrepancies. Specifically, it leverages Kullback-Leibler divergence to dynamically govern the modality transition from training to deployment, while incorporating representation alignment to facilitate a smooth shift from privileged states to RGB-D observations, thereby effectively resolving training-testing modality inconsistencies. Empirical evaluations demonstrate that the proposed approach improves the average test success rate from 0.71 to 0.93, yields substantial gains in budget-normalized area under the curve (AUC), and enhances overall learning efficiency by fivefold.
📝 Abstract
Vision-based reinforcement learning for robotic manipulation is sample-inefficient because RGB-D observations are high-dimensional and noisy. Privileged state information available in simulation can accelerate training, but its absence at test time creates a train-test modality gap. We propose FOCUS, a single-stage PPO framework that trains the critic on privileged state while automatically regulating whether the actor collects rollouts from RGB-D or privileged-state latents. Regulation is driven by the KL divergence between the action distributions induced by the two modalities, while representation alignment encourages consistent action selection across them. Together, these mechanisms limit RGB-D rollouts when the actor's action distributions from RGB-D and privileged-state latents disagree. As they align, RGB-D exposure increases, shifting on-policy training toward the RGB-D inputs used at test time. Across five manipulation tasks, FOCUS raises average test success from 0.71 to 0.93 relative to the strongest RGB-D-at-test baseline on each task. When accounting for each method's complete training pipeline, budget-normalized training-success AUC increases from 0.47 to 0.65. On Pick-and-Place, test success rises from 0.47 to 0.86, while AUC increases from 0.12 to 0.61, a 5.0x improvement in learning efficiency over the fixed interaction budget.
Problem

Research questions and friction points this paper is trying to address.

reinforcement learning
robotic manipulation
sample efficiency
modality gap
privileged states
Innovation

Methods, ideas, or system contributions that make the work stand out.

Privileged State
Modality Switching
Representation Alignment
Reinforcement Learning
Robotic Manipulation
🔎 Similar Papers
No similar papers found.