🤖 AI Summary
This work proposes UniCon, a unified framework that addresses the inefficiency of small-batch stochastic optimization commonly used in contrastive learning. By reformulating the contrastive alignment problem into an analytically solvable form, UniCon introduces a contrastive similarity weighting matrix and derives a closed-form global solution in a reproducing kernel Hilbert space (RKHS), thereby eliminating the need for conventional backpropagation. The approach seamlessly accommodates both linear and nonlinear encoders and supports diverse alignment paradigms, while also uncovering a fundamental connection between contrastive learning and spectral methods. Empirical evaluations demonstrate that UniCon substantially improves training efficiency across synthetic, unimodal, multimodal, and zero-shot tasks without compromising—indeed, often enhancing—generalization performance.
📝 Abstract
Contrastive objectives power state-of-the-art multimodal models, but their training remains slow, relying on long stochastic optimization. We propose a Unified Framework for Efficient Contrastive Alignment via Kernels (UniCon), which spans linear and nonlinear encoders as well as one-to-one and many-to-many alignments. At its core, UniCon introduces the contrastive similarity weight matrix $S(γ)$, which enables closed-form global solutions that provably replace minibatch back-propagation with exact updates. Through the lens of reproducing kernel Hilbert spaces (RKHS), UniCon provides a kernelized perspective that unifies contrastive alignment and reveals its connection to spectral methods. To validate the theory, we conduct experiments on synthetic, unimodal, multimodal, and zero-shot tasks, demonstrating that UniCon achieves substantial efficiency gains while preserving generality and strong empirical performance.