Wavelet Phase Diffusion for Structurally and Semantically Consistent Sim-to-Real Translation

📅 2026-07-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of large appearance discrepancies and difficulty in preserving structural and semantic consistency in simulation-to-real (Sim-to-Real) image translation by proposing a training-free, instance-level editing method. It introduces, for the first time, phase-preserving diffusion in the dual-tree complex wavelet packet transform domain, leveraging spatially localized phase injection and low-frequency randomization to achieve high-fidelity translation while maintaining structural and semantic coherence—without suffering from global spectral coupling artifacts typical of Fourier-domain approaches. Requiring only unpaired, open-domain data, the method significantly outperforms existing techniques on the vKITTI→KITTI benchmark, achieves near-paired-method realism in CARLA video translation, and reduces the ADE and FDE of a vision-language model (VLM) planner by 5.4% and 5.1%, respectively.
📝 Abstract
Simulation-to-reality translation must bridge the appearance gap between synthetic and real domains while preserving structural and semantic consistency. Conditioning-based methods achieve spatial alignment but introduce computationally expensive control modules. Paired-data methods achieve realism but rely on complex synthesis pipelines, often altering scene geometry and semantics. Training-free editing methods avoid both constraints but lack a learned appearance prior, limiting their perceptual quality. Recently proposed phase-preserving diffusion presents a promising alternative, but Fourier-domain formulations are constrained by global spectral coupling. This coupling induces spatial artifacts such as ringing and boundary leakage, thereby degrading structural and semantic consistency. We introduce Wavelet Phase Diffusion, which addresses this through two components. First, we operate in the Dual-Tree Complex Wavelet Packet Transform domain, whose localized wavelet packets enable spatially adaptive phase injection without global spectral interference. Second, Low-Frequency Randomization (LFR) replaces the low-frequency packet, decoupling the model from the synthetic illumination prior and enabling in-distribution real-world appearance. Both components train on unpaired open-domain data, and introduce negligible inference overhead. The spatial locality further enables instance-level translation, where individual objects or regions are translated to photorealistic appearance independently while the surrounding scene remains untranslated. On vKITTI $\to$ KITTI image translation, ours outperforms prior methods in realism and semantic consistency while maintaining competitive structural alignment. For CARLA video translation, ours approaches the realism of paired-data methods while reducing VLM planner ADE and FDE by $5.4\%$ and $5.1\%$, respectively.
Problem

Research questions and friction points this paper is trying to address.

sim-to-real translation
structural consistency
semantic consistency
phase-preserving diffusion
spatial artifacts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Wavelet Phase Diffusion
Dual-Tree Complex Wavelet Packet Transform
Low-Frequency Randomization
Sim-to-Real Translation
Instance-level Editing
🔎 Similar Papers
No similar papers found.