non-contrastive vl pretraining

Design and implement pretraining objectives and training pipelines that learn joint vision–language representations using non‑contrastive methods; specifically, build multimodal encoders and loss functions that predict cross‑modal targets with stop‑gradient targets (eschewing negatives, temperature scaling, or momentum teachers) and apply per‑modality distributional regularization to stabilize and shape alignment.

non-contrastivevlpretraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.56
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work explores the design space of natively multimodal foundation models, addressing how to effectively integrate vision and language beyond conventional language modeling. Building upon the Transfusion framework, the authors propose a unified pretraining approach from scratch that jointly leverages next-token prediction and diffusion-based generation, augmented with a Representation Autoencoder (RAE) to unify visual representations for both understanding and generation. The study reveals the complementary nature and asymmetric scaling behavior of vision and language data—where vision benefits more substantially from increased data volume—and employs a Mixture-of-Experts (MoE) architecture to enable efficient modality specialization and model expansion. Experiments demonstrate that unified pretraining naturally induces world modeling capabilities and significantly enhances performance on downstream tasks, laying a foundation for truly integrated multimodal foundation models.

foundation modelsmodality asymmetrymultimodal pretraining

This work proposes NOVA, a non-contrastive learning framework for vision–language alignment that eliminates the need for large batch sizes, negative sampling, momentum encoders, or gradient clipping—common requirements in existing contrastive methods that hinder training efficiency and stability. NOVA directly predicts the embeddings of a frozen ClinicalBERT text encoder from augmented image views, significantly simplifying the training pipeline. To regularize the learned representation distribution, the method introduces Sketched Isotropic Gaussian Regularization (SIGReg), which requires only a single hyperparameter. When combined with a Vision Transformer trained from scratch, NOVA achieves state-of-the-art performance on three zero-shot chest X-ray classification benchmarks using the MIMIC-CXR dataset, demonstrating both superior accuracy and enhanced training stability compared to existing baselines.

contrastive learninghyperparameter tuningnegative sampling

This work systematically investigates the core mechanisms of modality interaction in multimodal pretraining, focusing on knowledge transfer, synergistic effects, and fusion timing. Through experiments on both synthetic and large-scale real-world datasets, it provides the first empirical evidence of asymmetric cross-modal knowledge flow and demonstrates that data complexity governs whether modalities exhibit synergy or competition. The study further validates that early unified fusion consistently outperforms late alignment. Leveraging an architecture featuring shared attention and normalization layers with modality-specific feedforward components, the proposed approach is evaluated on a 13.5B mixture-of-experts model trained on 2 trillion tokens, confirming its effectiveness. Additionally, the paper introduces a highly efficient pretraining strategy that achieves strong generative performance using only 5% of the typical computational budget.

early unificationknowledge flowmodality interaction

This work addresses the prevalent text-dominant bias in existing vision-language models (VLMs), where visual signals are treated merely as passive inputs, leading to the loss of fine-grained visual details and coarse-grained multimodal understanding. To overcome this limitation, we propose Youtu-VL, a novel framework that introduces the Vision-Language Unified Autoregressive Supervision (VLUAS) paradigm. VLUAS unifies visual and linguistic tokens into a single autoregressive prediction sequence, enabling visual tokens to serve as prediction targets rather than just contextual inputs. This approach breaks away from conventional text-centric training paradigms and supports a wide range of vision-centric tasks without task-specific customization. Extensive experiments demonstrate that Youtu-VL achieves competitive performance on both general multimodal benchmarks and vision-intensive tasks, significantly enhancing visual detail preservation and joint multimodal modeling capabilities.

fine-grained visual informationmultimodal comprehensiontext-dominant bias

Exploring Transferability of Multimodal Adversarial Samples for Vision-Language Pre-training Models with Contrastive Learning

Aug 24, 2023
YW
Youze Wang
🏛️ Hefei University of Technology | Tsinghua University

Vision-language pretraining (VLP) models exhibit insufficient adversarial robustness in image-text joint tasks, particularly under cross-modal joint perturbations. This work proposes the first gradient-based multimodal adversarial attack method grounded in contrastive learning, capable of simultaneously generating imperceptible adversarial images and texts. Crucially, it introduces a novel joint optimization framework that unifies cross-modal (image-text) and intra-modal contrastive losses, thereby substantially enhancing the transferability of adversarial examples across model architectures in black-box settings. Evaluated on image-text retrieval and visual entailment tasks, the method significantly outperforms both unimodal and state-of-the-art multimodal attacks: it achieves an average transfer success rate improvement of 27.6% over existing approaches, demonstrating superior cross-architecture generalization and efficacy in practical adversarial scenarios.

Adversarial PerturbationsRobustnessVLP Models

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing vision-language pretraining approaches, which predominantly rely on contrastive learning and struggle to produce high-quality visual features suitable for dense prediction tasks. The authors propose the first end-to-end, fully non-contrastive pretraining method that eliminates the need for negative samples, temperature scaling, momentum encoders, or teacher-student mechanisms. Their approach achieves stable large-scale semantic alignment through cross-modal target prediction (with gradient stopping), intra-modal distribution regularization, and joint training. Under a frozen backbone setting, the method achieves state-of-the-art performance on GQA, VQAv2, and POPE, while significantly outperforming contrastive baselines on dense prediction tasks such as semantic segmentation, all without compromising global semantic understanding.

dense predictionfrozen backbonenon-contrastive learning

研究探讨了仅训练投影器而非整个多模态大语言模型骨干以适应新模ality的方法,证明此方法有效且能保持原有能力,同时提高训练效率。

fine-tuningmultimodal large language modelprojector

This work addresses the limitations of existing audio-visual self-supervised learning methods, which rely on modality-specific encoders and complex objective functions that hinder effective cross-modal synergy. The authors propose the first Joint Embedding Predictive Architecture (JEPA) for audio-visual representation learning, featuring a modality-agnostic unified encoder and a single predictive objective that jointly models intra- and inter-modal relationships to enable complementary information exchange across modalities. With a frozen ViT-g backbone, the method surpasses the previous best frozen baseline by 6.8 mAP on AudioSet-20K and outperforms fully fine-tuned models on ESC-50 and FSD50K. Notably, it achieves competitive performance on video tasks using only one-tenth of the video data, demonstrating substantially improved representation efficiency and generalization capability.

audio-visual learningcross-modal representationjoint embedding

This study addresses the unclear limits and scaling laws governing visual capability injection via lightweight projectors when large language model (LLM) weights remain frozen. Employing GLM—a text-only LLM—as the base model and Kimi as the vision encoder, this work proposes a reproducible training recipe utilizing a frontier-scale visual adapter comprising only 50M parameters. It systematically investigates how multimodal capabilities scale with LLM size under frozen pretrained weights. The primary contribution lies in successfully endowing a purely linguistic model with visual functionality while revealing the boundaries of such capabilities at scale. Comprehensive evaluations on the MMMU-Pro and BLINK benchmarks validate both overall performance and task-specific visual proficiency, thereby establishing a new paradigm for efficiently constructing multimodal large language models.

Frontier ScalesLanguage Model ScalingMultimodal Learning

This work addresses the lack of a systematic understanding of when cross-modal alignment (CA) and cross-modal prediction (CP) are effective in multimodal learning—a gap that often leads to suboptimal performance or even degradation relative to unimodal baselines. The authors propose a unified linear analytical framework under a structured signal–noise model with correlated interference, revealing complementary failure mechanisms of CA and CP. They introduce the first multimodal “phase diagram,” which delineates four distinct regimes: both methods succeed, only alignment works, only prediction works, or neither is effective. Leveraging separation ratio analysis, a unidirectional whitening mechanism, and a few-shot label-guided localization algorithm, this phase diagram enables practical guidance for method selection on real-world data. Experiments across synthetic, stereo vision, image–text, and astrophysical datasets validate its efficacy in identifying harmful multimodal configurations, offering a diagnostic tool for practitioners prior to model deployment.

cross-modal alignmentcross-modal predictionmultimodal learning

Hot Scholars

KS

Karan Singhal

Health AI at OpenAI
Health AIAI SafetyBeneficial AI
JB

Jelena Bratulić

PhD Student at University of Freiburg
Computer VisionDeep LearningMachine Learning
AV

Abhinav Valada

Professor & Director of Robot Learning Lab, University of Freiburg
RoboticsMachine LearningComputer VisionArtificial Intelligence