🤖 AI Summary
This study addresses the challenge of aligning Schrödinger bridges with human preferences and physical constraints during domain translation. To this end, we propose TSBM, a method that fine-tunes pretrained bridges toward reward-tilted objectives while strictly preserving the source distribution. The core innovation lies in introducing the concept of reward tilting from diffusion models into Schrödinger bridges for the first time, establishing an alternating optimization framework grounded in adjoint matching to provide rigorous theoretical support for post-training fine-tuning. Experimental results demonstrate the effectiveness of our approach on unpaired image translation tasks involving digit attributes in MNIST and facial features in CelebA.
📝 Abstract
Schrödinger bridges provide an entropy-regularized framework and a principled solution for unpaired domain translation. In practice, a pretrained bridge may need to be adapted to human preferences or physical constraints through a reward a problem closely related to reward tilting in diffusion models but underexplored for Schrödinger bridges. We introduce Tilted Schrödinger Bridge Matching (TSBM), a post-training method for fine-tuning a learned bridge $P$ between source $p_0$ and target $p_1$ toward a reward-tilted target $p_1^r\propto p_1e^r$, while preserving source $p_0$. We formulate this adaptation as alternating optimization initialized from $P$, provide theoretical justification, and derive a practical algorithm based on Adjoint Matching. We evaluate TSBM on unpaired image-to-image translation targeting digit properties in MNIST and facial attributes in CelebA.