The Domain Is a Residue: Adapting Self-Supervised Features, Not Generators

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the persistent challenge of residual source-domain style artifacts in generators for unpaired image-to-image translation. To this end, we propose the Representation Feature Adapter (RFA), which models domain discrepancies as residuals within the DINO self-supervised feature space. By training a lightweight adapter to translate these style residuals, RFA achieves effective domain adaptation while keeping both the encoder and decoder frozen throughout the process. Compared with conventional generator fine-tuning strategies, our method demonstrates superior performance on dehazing and simulation-to-real tasks, outperforming existing baselines while reducing parameter count by 160 times and substantially shortening training time. Ultimately, RFA enables efficient style disentanglement alongside faithful preservation of scene structure.
📝 Abstract
Clearing fog, rain or snow from footage, or turning renders into photographs, must remove the source domain and keep the scene. Unpaired translators carry it through because their generator sees the source appearance (pixels, a near-invertible latent or a control map) and keeps it. A DINO feature map fixes what is in the scene and carries weather, lighting and rendering style as a residue of 13 to 14% of the feature norm. We propose the Representation Feature Adapter (RFA), a 2.9M-parameter network that moves this residue. We train only the adapter and its discriminators; the encoder and a feature-conditioned decoder, trained once for all conditions, stay frozen. Against CycleGAN-Turbo it is ahead on both metrics on fog and on KID on night, and level within noise on snow, rain and haze. On sim-to-real it leads REGEN and HyPER-GAN on both metrics. Only the RFA removes the rain while keeping the scene. The removal costs scene structure: CycleGAN-Turbo keeps more on every condition but fog. On VAE latents the identical adapter collapses to the identity, and decoders from other groups that never saw it render its output. The RFA has about 160 times fewer trainable parameters than CycleGAN-Turbo and under a fifth of its per-condition training time.
Problem

Research questions and friction points this paper is trying to address.

unpaired image translation
domain adaptation
self-supervised features
scene preservation
sim-to-real
Innovation

Methods, ideas, or system contributions that make the work stand out.

Representation Feature Adapter
Self-Supervised Features
Domain Residue
Unpaired Image Translation
Parameter Efficiency