unpaired image translation

Design and implement models and algorithms that translate images between visual domains without paired input-output examples, covering domain translation, asymmetric and domain-adaptive style transfer, and approaches based on stochastic transport such as neural Schrödinger bridge synthesis. This work builds mappings or entropy-regularized transport processes that preserve spatial continuity and semantic consistency (including synthetic pixel-level annotations), reduce cross-domain distribution shift, and can operate using only unlabeled real reference images when required.

unpairedimagetranslation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.39
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$193K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

CycleDiff: Cycle Diffusion Models for Unpaired Image-to-image Translation

Aug 08, 2025
SZ
Shilong Zou
🏛️ National University of Defense Technology

To address unpaired cross-domain image translation, this paper proposes a diffusion-based cycle learning framework that jointly optimizes diffusion denoising and image translation to mitigate the local optima problem inherent in conventional approaches. Its key contributions are: (1) a time-dependent translation network that dynamically aligns diffusion timesteps with domain mappings; and (2) a diffusion-based clean-signal extraction mechanism that end-to-end disentangles structural and textural components. Integrated with cycle-consistency constraints and unpaired adversarial training, the model achieves state-of-the-art performance on bidirectional multimodal translation tasks—including RGB ↔ edge, semantic, and depth domains—yielding outputs with both high fidelity and strong structural consistency.

Aligning diffusion and translation processes for global optimizationImproving cross-domain translation fidelity and structural consistencyUnpaired image translation without paired training data

This work addresses the limitations of existing generative semantic communication methods, which rely on Gaussian priors and consequently suffer from severe hallucinations and high computational overhead in narrowband, high-noise channels. To overcome these issues, the authors propose a Schrödinger Bridge-based Generative Semantic Communication (SBGSC) framework that dispenses with Gaussian assumptions by constructing an optimal transport trajectory between semantic and image distributions, enabling direct generative decoding. By reformulating the nonlinear drift term of diffusion models and introducing a self-consistent guidance strategy for non-Markovian generation, SBGSC effectively learns the underlying velocity field, drastically reducing sampling steps while suppressing hallucinatory artifacts. Experimental results demonstrate that SBGSC improves the Fréchet Inception Distance (FID) by at least 38%, increases Structural Similarity Index (SSIM) by 49.3%, and accelerates inference by over eightfold compared to current state-of-the-art methods.

Computational CostGenerative Semantic CommunicationHallucination

Implicit Image-to-Image Schrodinger Bridge for Image Restoration

Mar 10, 2024
YW
Yuang Wang
🏛️ Tsinghua University | Harvard Medical School | Massachusetts General Hospital | Zhejiang University

Diffusion-based image restoration suffers from slow inference; although I²SB improves efficiency, further acceleration is needed. This paper proposes Implicit Image-to-Image Schrödinger Bridge (I³SB), the first non-Markovian generative framework for restoration: at each step, the initial degraded image is explicitly injected as an implicit condition, enabling zero-shot adaptation of pre-trained I²SB models without retraining while preserving marginal distribution consistency. I³SB integrates Schrödinger bridge theory, score-based modeling, and implicit conditional injection to jointly leverage score-based priors and degradation guidance. Evaluated on diverse degradation tasks—including medical, facial, and natural images—I³SB achieves perceptual quality comparable to I²SB using significantly fewer sampling steps, while substantially improving reconstruction realism and structural fidelity.

Accelerate image restoration using implicit Schrodinger BridgeApply non-Markovian framework to preserve corrupted image informationReduce generative steps while maintaining perceptual quality

A Diffusion Model Translator for Efficient Image-to-Image Translation

Jul 30, 2024
MX
Mengfei Xia
🏛️ Tsinghua University | Shanghai Jiao Tong University | Texas A&M University

Existing image-to-image (I2I) translation diffusion models redundantly inject the source image at every denoising step, leading to inefficient inference. This work proposes the lightweight Diffusion Model Translator (DMT), which—through the first theoretical analysis—demonstrates that a single inter-domain distribution transfer suffices for high-fidelity I2I translation. Guided by this insight, DMT introduces a compact translation module that performs distribution alignment only at a critical intermediate timestep. Built upon the DDPM framework, DMT integrates probabilistic distribution shift analysis, an adaptive optimal timestep selection strategy, and a streamlined architecture. Extensive experiments on style transfer, colorization, semantic segmentation map generation, and sketch-to-color tasks show that DMT achieves state-of-the-art performance with significantly faster inference—averaging 3.2× speedup—while simultaneously delivering superior image quality.

De-noisingImage-to-Image TranslationSpeed Optimization

Exploring the Design Space of Diffusion Bridge Models via Stochasticity Control

Oct 28, 2024
SZ
Shaorong Zhang
🏛️ University of California Riverside

Existing diffusion bridge and stochastic interpolation models for pixel-space image-to-image translation suffer from technical fragmentation due to incompatible mathematical assumptions and neglect the insufficient diversity under fixed sampling budgets. This paper proposes the Stochasticity-Controlled Diffusion Bridge (SDB), the first framework to jointly regulate three sources of stochasticity—sampling SDE dynamics, transition kernels, and base distributions—along the noise-source dimension. SDB avoids training and sampling singularities and introduces a differentiable diversity metric. Built upon extended diffusion bridge theory and SDE-based modeling, SDB integrates controllable noise injection and joint FID/diversity evaluation. Empirically, SDB achieves state-of-the-art performance: it maintains high visual fidelity while accelerating sampling by 5× over baselines, significantly reducing FID, and substantially improving generation diversity.

Addressing low sample diversity in fixed conditionsEnhancing sampling efficiency and image qualityUnifying diffusion bridge models for image translation

Latest Papers

What's happening recently
View more

This work addresses the challenge of weakly aligned training data in image-to-image translation, which arises from discrepancies in acquisition conditions such as asynchronous capture, illumination variations, or registration errors. The authors propose an alignment-aware bridge matching method that explicitly models alignment information as a conditional variable within a bridge matching framework. By introducing an alignment score to distinguish genuine semantic correspondences from misalignment artifacts, the method enables fidelity-controllable translation during inference. Built upon a unified framework bridging bridge matching and flow matching, the approach effectively leverages weakly aligned data and significantly outperforms existing baselines—including GANs, diffusion models, and Schrödinger bridge methods—on tasks such as cross-sensor super-resolution and unsupervised domain adaptation.

alignment artifactsdomain adaptationimage-to-image translation

This study addresses the challenge of aligning Schrödinger bridges with human preferences and physical constraints during domain translation. To this end, we propose TSBM, a method that fine-tunes pretrained bridges toward reward-tilted objectives while strictly preserving the source distribution. The core innovation lies in introducing the concept of reward tilting from diffusion models into Schrödinger bridges for the first time, establishing an alternating optimization framework grounded in adjoint matching to provide rigorous theoretical support for post-training fine-tuning. Experimental results demonstrate the effectiveness of our approach on unpaired image translation tasks involving digit attributes in MNIST and facial features in CelebA.

human preferencespost-training adaptationreward tilting

This study addresses the lack of spatial structure in source distributions of flow matching models, which violates the inductive bias of image locality. We propose StructFlow, a method that encodes spatial locality into source noise to construct a structured source distribution, thereby achieving geometric alignment between transport paths and image regions. Furthermore, it incorporates lightweight post-training fine-tuning to adapt to Diffusion Transformer architectures. Experimental results demonstrate that StructFlow achieves competitive generation quality while significantly improving performance in locally controllable resynthesis, structure preservation, and semantic interpolation. Consequently, this approach effectively resolves persistent challenges regarding local editing and structural consistency in generative modeling.

Flow MatchingImage GenerationInductive Bias

Hot Scholars

DC

Daniel Cremers

Technical University of Munich
Computer VisionMachine LearningOptimizationRobotics
CC

Ching-Chun Huang

National Yang Ming Chiao Tung University
Computer VisionSignal ProcessingMachine Learning
AS

Abhishek Saroha

Unknown affiliation
Machine LearningComputer Vision3D Reconstruction
QH

Qiyuan He

City University of Hong Kong
Semiconductor interfaces