Score
Design and implement pretraining objectives and training pipelines that learn joint vision–language representations using non‑contrastive methods; specifically, build multimodal encoders and loss functions that predict cross‑modal targets with stop‑gradient targets (eschewing negatives, temperature scaling, or momentum teachers) and apply per‑modality distributional regularization to stabilize and shape alignment.
This work explores the design space of natively multimodal foundation models, addressing how to effectively integrate vision and language beyond conventional language modeling. Building upon the Transfusion framework, the authors propose a unified pretraining approach from scratch that jointly leverages next-token prediction and diffusion-based generation, augmented with a Representation Autoencoder (RAE) to unify visual representations for both understanding and generation. The study reveals the complementary nature and asymmetric scaling behavior of vision and language data—where vision benefits more substantially from increased data volume—and employs a Mixture-of-Experts (MoE) architecture to enable efficient modality specialization and model expansion. Experiments demonstrate that unified pretraining naturally induces world modeling capabilities and significantly enhances performance on downstream tasks, laying a foundation for truly integrated multimodal foundation models.
This work proposes NOVA, a non-contrastive learning framework for vision–language alignment that eliminates the need for large batch sizes, negative sampling, momentum encoders, or gradient clipping—common requirements in existing contrastive methods that hinder training efficiency and stability. NOVA directly predicts the embeddings of a frozen ClinicalBERT text encoder from augmented image views, significantly simplifying the training pipeline. To regularize the learned representation distribution, the method introduces Sketched Isotropic Gaussian Regularization (SIGReg), which requires only a single hyperparameter. When combined with a Vision Transformer trained from scratch, NOVA achieves state-of-the-art performance on three zero-shot chest X-ray classification benchmarks using the MIMIC-CXR dataset, demonstrating both superior accuracy and enhanced training stability compared to existing baselines.
This work systematically investigates the core mechanisms of modality interaction in multimodal pretraining, focusing on knowledge transfer, synergistic effects, and fusion timing. Through experiments on both synthetic and large-scale real-world datasets, it provides the first empirical evidence of asymmetric cross-modal knowledge flow and demonstrates that data complexity governs whether modalities exhibit synergy or competition. The study further validates that early unified fusion consistently outperforms late alignment. Leveraging an architecture featuring shared attention and normalization layers with modality-specific feedforward components, the proposed approach is evaluated on a 13.5B mixture-of-experts model trained on 2 trillion tokens, confirming its effectiveness. Additionally, the paper introduces a highly efficient pretraining strategy that achieves strong generative performance using only 5% of the typical computational budget.
This work addresses the prevalent text-dominant bias in existing vision-language models (VLMs), where visual signals are treated merely as passive inputs, leading to the loss of fine-grained visual details and coarse-grained multimodal understanding. To overcome this limitation, we propose Youtu-VL, a novel framework that introduces the Vision-Language Unified Autoregressive Supervision (VLUAS) paradigm. VLUAS unifies visual and linguistic tokens into a single autoregressive prediction sequence, enabling visual tokens to serve as prediction targets rather than just contextual inputs. This approach breaks away from conventional text-centric training paradigms and supports a wide range of vision-centric tasks without task-specific customization. Extensive experiments demonstrate that Youtu-VL achieves competitive performance on both general multimodal benchmarks and vision-intensive tasks, significantly enhancing visual detail preservation and joint multimodal modeling capabilities.
Vision-language pretraining (VLP) models exhibit insufficient adversarial robustness in image-text joint tasks, particularly under cross-modal joint perturbations. This work proposes the first gradient-based multimodal adversarial attack method grounded in contrastive learning, capable of simultaneously generating imperceptible adversarial images and texts. Crucially, it introduces a novel joint optimization framework that unifies cross-modal (image-text) and intra-modal contrastive losses, thereby substantially enhancing the transferability of adversarial examples across model architectures in black-box settings. Evaluated on image-text retrieval and visual entailment tasks, the method significantly outperforms both unimodal and state-of-the-art multimodal attacks: it achieves an average transfer success rate improvement of 27.6% over existing approaches, demonstrating superior cross-architecture generalization and efficacy in practical adversarial scenarios.
This work addresses the limitations of existing vision-language pretraining approaches, which predominantly rely on contrastive learning and struggle to produce high-quality visual features suitable for dense prediction tasks. The authors propose the first end-to-end, fully non-contrastive pretraining method that eliminates the need for negative samples, temperature scaling, momentum encoders, or teacher-student mechanisms. Their approach achieves stable large-scale semantic alignment through cross-modal target prediction (with gradient stopping), intra-modal distribution regularization, and joint training. Under a frozen backbone setting, the method achieves state-of-the-art performance on GQA, VQAv2, and POPE, while significantly outperforming contrastive baselines on dense prediction tasks such as semantic segmentation, all without compromising global semantic understanding.
研究探讨了仅训练投影器而非整个多模态大语言模型骨干以适应新模ality的方法,证明此方法有效且能保持原有能力,同时提高训练效率。
This work addresses the limitations of existing audio-visual self-supervised learning methods, which rely on modality-specific encoders and complex objective functions that hinder effective cross-modal synergy. The authors propose the first Joint Embedding Predictive Architecture (JEPA) for audio-visual representation learning, featuring a modality-agnostic unified encoder and a single predictive objective that jointly models intra- and inter-modal relationships to enable complementary information exchange across modalities. With a frozen ViT-g backbone, the method surpasses the previous best frozen baseline by 6.8 mAP on AudioSet-20K and outperforms fully fine-tuned models on ESC-50 and FSD50K. Notably, it achieves competitive performance on video tasks using only one-tenth of the video data, demonstrating substantially improved representation efficiency and generalization capability.
This study addresses the unclear limits and scaling laws governing visual capability injection via lightweight projectors when large language model (LLM) weights remain frozen. Employing GLM—a text-only LLM—as the base model and Kimi as the vision encoder, this work proposes a reproducible training recipe utilizing a frontier-scale visual adapter comprising only 50M parameters. It systematically investigates how multimodal capabilities scale with LLM size under frozen pretrained weights. The primary contribution lies in successfully endowing a purely linguistic model with visual functionality while revealing the boundaries of such capabilities at scale. Comprehensive evaluations on the MMMU-Pro and BLINK benchmarks validate both overall performance and task-specific visual proficiency, thereby establishing a new paradigm for efficiently constructing multimodal large language models.
This work addresses the lack of a systematic understanding of when cross-modal alignment (CA) and cross-modal prediction (CP) are effective in multimodal learning—a gap that often leads to suboptimal performance or even degradation relative to unimodal baselines. The authors propose a unified linear analytical framework under a structured signal–noise model with correlated interference, revealing complementary failure mechanisms of CA and CP. They introduce the first multimodal “phase diagram,” which delineates four distinct regimes: both methods succeed, only alignment works, only prediction works, or neither is effective. Leveraging separation ratio analysis, a unidirectional whitening mechanism, and a few-shot label-guided localization algorithm, this phase diagram enables practical guidance for method selection on real-world data. Experiments across synthetic, stereo vision, image–text, and astrophysical datasets validate its efficacy in identifying harmful multimodal configurations, offering a diagnostic tool for practitioners prior to model deployment.