Digital Twin-Driven Real2Sim2Real: Simulator-Conditioned Generation via Paired Driving-Scene Reconstruction

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance of autonomous driving perception on costly annotations and the inherent trade-off between visual realism and label-free generation in synthetic data augmentation. To overcome these limitations, this work proposes a digital twin-driven Real2Sim2Real framework that pioneers georeferenced digital twin-based simulator-conditioned diffusion generation. By integrating scene reconstruction with geographic registration, the method synthesizes high-fidelity driving images that maintain precise visual alignment with real-world environments while preserving cost-free annotations. Experimental results demonstrate that the proposed DET R3D achieves 93.18% mAP relative to the real-data baseline. Furthermore, joint training with synthetic and real data surpasses the performance ceiling of purely real-data training, significantly reducing data acquisition costs.
📝 Abstract
Camera-based 3D perception for autonomous driving relies heavily on large annotated datasets, and deploying such a system to a new target region typically requires data collection and annotation. Generative augmentation has been proposed to reduce this cost, but existing approaches face a fundamental trade-off: label-conditioned methods consume the very annotations they aim to replace, while simulator-conditioned methods offer free annotations but lack visual grounding to specific real environments. This work investigates the extent to which a digital-twin-driven Real2Sim2Real pipeline (DT-R2S2R) can substitute for target-region real data. By reconstructing recorded driving clips inside a georeferenced digital twin (DT-R2S), we condition a diffusion model on geometrically aligned simulator renderings, establishing a digital twin-grounded Sim2Real model (DT-S2R). As a result, DT-S2R synthesizes photorealistic driving images given low-cost yet georeferenced simulator data across both reconstructed and novel simulator scenes within digital-twin coverage. The efficacy of generated data is verified on diverse 3D detectors. DETR3D, especially, reports 93.18% of mAP obtained by a target-region real-data oracle, without employing target images for detector training. Furthermore, simple co-training with existing out-of-target real data outperforms the oracle. Thus, DT-R2S2R can substantially reduce the cost of manual on-site data collection and annotation in digital twin-available districts, providing a practical foundation for scaling 3D perception.
Problem

Research questions and friction points this paper is trying to address.

Autonomous Driving
3D Perception
Data Augmentation
Digital Twin
Real2Sim2Real
Innovation

Methods, ideas, or system contributions that make the work stand out.

Digital Twin
Real2Sim2Real
Diffusion Model
3D Perception
Autonomous Driving
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hojun Lim
Technical University of Munich, Germany; School of Engineering & Design, Department of Mobility Systems Engineering, Institute of Automotive Technology and Munich Institute of Robotics and Machine Intelligence (MIRMI)
Hyeongseok Jeon
Hyeongseok Jeon
MORAI Inc., Seoul, South Korea
D
Donghyun Kim
AVP Division, Hyundai Motor Company, Gyeonggi-do, South Korea
S
Soonyoung Jung
AVP Division, Hyundai Motor Company, Gyeonggi-do, South Korea
Heecheol Yoo
Heecheol Yoo
MORAI Inc.
Computer VisionAutonomous Driving