RoboRender: Robot-Oriented Video Generation for Visual Sim-to-Real Transfer

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of transferring simulation-trained policies to the real world due to visual discrepancies. To this end, we propose a robot-oriented video generation framework that translates simulated trajectories into photorealistic RGB videos for policy learning, synthesizing authentic textured backgrounds while preserving geometric and kinematic information. The core innovation lies in conditioning the video generation model on depth videos, language instructions, and robot masks, enabling zero-shot deployment without intermediate representations or additional perception modules. Experimental results demonstrate that the proposed method achieves an average success rate of 71% on real-world tasks, significantly outperforming both raw simulation rendering and conventional domain randomization approaches.
📝 Abstract
Simulation enables large-scale, low-cost robot data generation, but policies trained in simulation often fail to transfer to the real world due to the sim-to-real visual discrepancies. Existing approaches often rely on intermediate representations, which can discard rich semantic information or require additional perception modules at deployment. We address this visual sim-to-real gap with RoboRender, a framework that converts simulated trajectories into photorealistic RGB videos for policy learning. RoboRender trains a robot-oriented video generation model conditioned on simulated depth videos, language instructions, and robot RGB mask videos, preserving simulator geometry, robot motion, and action labels while synthesizing realistic textures, backgrounds, and distractors. The generated RGB videos are paired with simulator-provided states and actions to train policies for zero-shot real-world deployment. On robot video test sets, our video model outperforms depth-conditioned video generation baselines in generation quality. In real-world experiments across pick-and-place, articulated-object manipulation, and mobile manipulation tasks, policies trained on RoboRender-generated data achieve a 71% average success rate, outperforming raw simulation renderings and conventional visual domain randomization by approximately 7.1x and 3.6x, respectively. We further show that policy performance improves with more generated videos per simulation trajectory, increasing opening-task success by 65 percentage points. These results demonstrate that generative video rendering mitigates the visual sim-to-real gap for zero-shot policy transfer. Project website: https://robo-render.github.io/.
Problem

Research questions and friction points this paper is trying to address.

sim-to-real gap
visual sim-to-real transfer
robot policy learning
video generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sim-to-Real Transfer
Video Generation
Robot Manipulation
Zero-Shot Policy Transfer
Domain Randomization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.