KALE: Kernel Alignment with Loss Equilibration for Stable CLIP-DINOv2 Alignment at Web Scale

📅 2026-07-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of ineffective alignment signals in existing CLIP and DINOv2 methods when trained on web-scale noisy datasets like CC12M, where fixed loss weights hinder visual representation learning. To overcome this, the authors propose KALE—the first adaptive alignment mechanism tailored for large-scale noisy data—which dynamically adjusts the CLIP–DINOv2 alignment loss weight to preserve effective gradient signals without requiring dataset-specific hyperparameter tuning. Integrating a kernel alignment framework, an adaptive loss balancing controller, and an aggressive learning rate decay strategy, KALE achieves a +2.00 average improvement in zero-shot performance across 11 standard benchmarks on a 3.3M-image subset, substantially outperforming KUEA (+1.29) and delivering consistent gains in SVHN linear probing and image–text retrieval tasks.
📝 Abstract
Kernel-based alignment of CLIP toward a vision centric teacher such as DINOv2 (KUEA) improves CLIP's visual representations while preserving text-encoder compatibility, using a fixed trade-off weight tuned on curated ImageNet-1K. We ask whether this transfers to noisy, web-scale data (CC12M) and find that it does not: the alignment term's weighted contribution falls to about 0.2% of the clean term, so under any fixed weight its gradient is effectively inert. We introduce KALE, a loss-equilibration controller that tracks both losses and adaptively rescales the alignment weight toward a target ratio, restoring the signal with no per-dataset tuning; reaching balance requires increasing the weight by roughly four orders of magnitude, and the required value is configuration-dependent, so no fixed scalar suffices. We characterize the resulting regime: a bounded high learning rate and a decaying schedule with a moderate floor are needed for stability, and the controller equilibrates rather than diverging. On a 3.3M-image CC12M subset, the aligned model preserves image-text retrieval and reproducibly improves SVHN linear probing; zero-shot improves by +2.00 over CLIP on the standard 11-dataset average, exceeding KUEA's +1.29. We report all results with explicit run-to-run variance and base our conclusions on the metrics that are stable across runs.
Problem

Research questions and friction points this paper is trying to address.

CLIP-DINOv2 alignment
web-scale data
loss equilibration
noisy data
kernel alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

loss equilibration
CLIP-DINOv2 alignment
web-scale training
adaptive weighting
kernel alignment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Michał Pawłowicz
Meta AI