Score
Design and implement models and algorithms that translate images between visual domains without paired input-output examples, covering domain translation, asymmetric and domain-adaptive style transfer, and approaches based on stochastic transport such as neural Schrödinger bridge synthesis. This work builds mappings or entropy-regularized transport processes that preserve spatial continuity and semantic consistency (including synthetic pixel-level annotations), reduce cross-domain distribution shift, and can operate using only unlabeled real reference images when required.
To address unpaired cross-domain image translation, this paper proposes a diffusion-based cycle learning framework that jointly optimizes diffusion denoising and image translation to mitigate the local optima problem inherent in conventional approaches. Its key contributions are: (1) a time-dependent translation network that dynamically aligns diffusion timesteps with domain mappings; and (2) a diffusion-based clean-signal extraction mechanism that end-to-end disentangles structural and textural components. Integrated with cycle-consistency constraints and unpaired adversarial training, the model achieves state-of-the-art performance on bidirectional multimodal translation tasks—including RGB ↔ edge, semantic, and depth domains—yielding outputs with both high fidelity and strong structural consistency.
This work addresses the limitations of existing generative semantic communication methods, which rely on Gaussian priors and consequently suffer from severe hallucinations and high computational overhead in narrowband, high-noise channels. To overcome these issues, the authors propose a Schrödinger Bridge-based Generative Semantic Communication (SBGSC) framework that dispenses with Gaussian assumptions by constructing an optimal transport trajectory between semantic and image distributions, enabling direct generative decoding. By reformulating the nonlinear drift term of diffusion models and introducing a self-consistent guidance strategy for non-Markovian generation, SBGSC effectively learns the underlying velocity field, drastically reducing sampling steps while suppressing hallucinatory artifacts. Experimental results demonstrate that SBGSC improves the Fréchet Inception Distance (FID) by at least 38%, increases Structural Similarity Index (SSIM) by 49.3%, and accelerates inference by over eightfold compared to current state-of-the-art methods.
Diffusion-based image restoration suffers from slow inference; although I²SB improves efficiency, further acceleration is needed. This paper proposes Implicit Image-to-Image Schrödinger Bridge (I³SB), the first non-Markovian generative framework for restoration: at each step, the initial degraded image is explicitly injected as an implicit condition, enabling zero-shot adaptation of pre-trained I²SB models without retraining while preserving marginal distribution consistency. I³SB integrates Schrödinger bridge theory, score-based modeling, and implicit conditional injection to jointly leverage score-based priors and degradation guidance. Evaluated on diverse degradation tasks—including medical, facial, and natural images—I³SB achieves perceptual quality comparable to I²SB using significantly fewer sampling steps, while substantially improving reconstruction realism and structural fidelity.
Existing image-to-image (I2I) translation diffusion models redundantly inject the source image at every denoising step, leading to inefficient inference. This work proposes the lightweight Diffusion Model Translator (DMT), which—through the first theoretical analysis—demonstrates that a single inter-domain distribution transfer suffices for high-fidelity I2I translation. Guided by this insight, DMT introduces a compact translation module that performs distribution alignment only at a critical intermediate timestep. Built upon the DDPM framework, DMT integrates probabilistic distribution shift analysis, an adaptive optimal timestep selection strategy, and a streamlined architecture. Extensive experiments on style transfer, colorization, semantic segmentation map generation, and sketch-to-color tasks show that DMT achieves state-of-the-art performance with significantly faster inference—averaging 3.2× speedup—while simultaneously delivering superior image quality.
Existing diffusion bridge and stochastic interpolation models for pixel-space image-to-image translation suffer from technical fragmentation due to incompatible mathematical assumptions and neglect the insufficient diversity under fixed sampling budgets. This paper proposes the Stochasticity-Controlled Diffusion Bridge (SDB), the first framework to jointly regulate three sources of stochasticity—sampling SDE dynamics, transition kernels, and base distributions—along the noise-source dimension. SDB avoids training and sampling singularities and introduces a differentiable diversity metric. Built upon extended diffusion bridge theory and SDE-based modeling, SDB integrates controllable noise injection and joint FID/diversity evaluation. Empirically, SDB achieves state-of-the-art performance: it maintains high visual fidelity while accelerating sampling by 5× over baselines, significantly reducing FID, and substantially improving generation diversity.
This work addresses the challenge of weakly aligned training data in image-to-image translation, which arises from discrepancies in acquisition conditions such as asynchronous capture, illumination variations, or registration errors. The authors propose an alignment-aware bridge matching method that explicitly models alignment information as a conditional variable within a bridge matching framework. By introducing an alignment score to distinguish genuine semantic correspondences from misalignment artifacts, the method enables fidelity-controllable translation during inference. Built upon a unified framework bridging bridge matching and flow matching, the approach effectively leverages weakly aligned data and significantly outperforms existing baselines—including GANs, diffusion models, and Schrödinger bridge methods—on tasks such as cross-sensor super-resolution and unsupervised domain adaptation.
本文提出离散扩散桥(DDB)框架,通过混合吸收机制和信息引导噪声调度解决图像转换和生成中的时空错位问题。
This study addresses the challenge of aligning Schrödinger bridges with human preferences and physical constraints during domain translation. To this end, we propose TSBM, a method that fine-tunes pretrained bridges toward reward-tilted objectives while strictly preserving the source distribution. The core innovation lies in introducing the concept of reward tilting from diffusion models into Schrödinger bridges for the first time, establishing an alternating optimization framework grounded in adjoint matching to provide rigorous theoretical support for post-training fine-tuning. Experimental results demonstrate the effectiveness of our approach on unpaired image translation tasks involving digit attributes in MNIST and facial features in CelebA.
This study addresses the lack of spatial structure in source distributions of flow matching models, which violates the inductive bias of image locality. We propose StructFlow, a method that encodes spatial locality into source noise to construct a structured source distribution, thereby achieving geometric alignment between transport paths and image regions. Furthermore, it incorporates lightweight post-training fine-tuning to adapt to Diffusion Transformer architectures. Experimental results demonstrate that StructFlow achieves competitive generation quality while significantly improving performance in locally controllable resynthesis, structure preservation, and semantic interpolation. Consequently, this approach effectively resolves persistent challenges regarding local editing and structural consistency in generative modeling.
为解决多模态翻译的灵活性和双向性问题,提出BIT方法,通过随机微积分实现从文本到图像及反向的生成路径。