Score
Designs training curricula and optimization procedures that progressively incorporate multiple data modalities—using hierarchical or stage-wise training, modality-aware objectives, curriculum schedules, and KL-based routing supervision—to stabilize learning of unified multimodal token representations and support variable-length multimodal sequences. Builds adaptive loss-weighting and routing schemes to mitigate heterogeneous-information imbalance, improve joint generation and alignment, and stabilize generation quality across training stages.
This work addresses the parameter efficiency challenge in visual–language large models (VLLMs) for effective cross-modal fusion. We systematically analyze 34 state-of-the-art VLLMs and, for the first time, unify their training paradigms into three categories—single-stage fine-tuning, two-stage fine-tuning, and direct adaptation—establishing the first taxonomy of VLLM efficiency grounded in training methodology. Our study fills a critical gap by providing the first systematic analysis of direct adaptation, empirically demonstrating that it achieves over 90% of two-stage fine-tuning performance with less than 1% parameter overhead. We comprehensively examine core components—including LLM backbones, vision encoders, multimodal fusion architectures, parameter-efficient adaptation techniques (e.g., LoRA, Adapters), and evaluation protocols—and synthesize key benchmarks and metrics. The work delivers both a theoretical framework and empirical evidence to advance efficient multimodal modeling.
This study addresses the unclear trade-offs among diverse capabilities—such as general visual understanding, structured reasoning, and fine-grained OCR—in multimodal instruction tuning under mixed data regimes, particularly the lack of systematic investigation into how data organization influences these trade-offs. Treating data scheduling as a first-order design variable while holding model architecture and optimization settings fixed, the work compares four strategies: direct mixing, curriculum learning, balanced sampling, and reverse curriculum. Results demonstrate that curriculum-based training—sequencing tasks from general comprehension to specialized skills—achieves superior overall performance and structured reasoning while accelerating convergence. Balanced sampling improves OCR accuracy at the cost of capability imbalance, whereas reverse curriculum degrades performance and induces optimization instability. The findings highlight the critical role of training sequence in shaping the capability distribution of multimodal models.
Existing reinforcement learning post-training methods fail to differentiate among multimodal inputs, leading to high policy gradient variance, slow convergence, and poor robustness to missing modalities or distribution shifts. To address this, this work proposes the first task-level benchmark for minimal modality-combination annotation and introduces MAPO, a modality-aware policy optimization framework that integrates hierarchical batching, adaptive weighting, and a curriculum scheduling mechanism based on signal-combination difficulty. Experiments on MAPLE-bench demonstrate that the proposed approach reduces the accuracy gap between unimodal and multimodal settings by 30.24%, accelerates convergence by a factor of 3.18, and maintains stable performance across diverse modality-missing scenarios.
This work addresses the optimization instability in multimodal large language models during autoregressive training, which arises from gradient heterogeneity across modalities and hinders scaling to large batch sizes. The authors propose ML-FOP-SOAP, a novel framework that introduces the second-order preconditioning method SOAP into multimodal training for the first time. By integrating Fisher orthogonal projection to mitigate inter-modal competition and incorporating multi-level variance correction with hierarchical folding strategies, the approach enables efficient co-optimization of visual generation and textual understanding under low computational overhead. Evaluated on Janus and Emu3, ML-FOP-SOAP significantly enhances training stability and sample efficiency, supporting stable training with batch sizes up to 8192. It achieves a 1.4× improvement in sample efficiency and reduces training time by 1.5× compared to AdamW.
Existing large models struggle to jointly model and freely generate interleaved multimodal sequences—such as speech, text, images, and video—within a unified framework. This paper introduces MIO, an end-to-end autoregressive multimodal foundation model, and proposes the first four-stage causal multimodal training paradigm: alignment pretraining → interleaving pretraining → speech-augmented pretraining → multitask supervised fine-tuning. MIO employs discrete multimodal tokenization and unified causal modeling to enable true any-to-any cross-modal understanding and generation. It supports arbitrary input/output modality combinations and long-range interleaved sequence generation, unlocking novel capabilities including chain-of-visual-reasoning and instruction-driven image editing. On comprehensive cross-modal benchmarks, MIO matches or surpasses state-of-the-art dual-modal, any-to-any, and unimodal specialized models. Notably, it achieves significant performance gains on complex interleaved tasks—such as video-text generation and visual-guided instruction generation.
Traditional multimodal approaches rely on modality-specific encoders and late fusion, limiting scalability and cross-modal generalization. This work proposes a unified paradigm that reformulates diverse multimodal tasks—including text, image, audio, and video processing—as “next-frame prediction,” with all inputs and outputs represented as serialized video frames, enabling end-to-end, single-model inference. It pioneers the task-redefinition strategy—previously confined to NLP—within multimodal learning, eliminating modality-specific design in favor of a fully modality-agnostic architecture. The framework integrates cross-modal tokenization, sequential frame representation, autoregressive modeling, and a shared Transformer decoder, supporting joint multi-task pretraining. Empirical evaluation demonstrates strong zero-shot and few-shot generalization across text-to-text, image-to-text, video-to-video, video-to-text, and audio-to-text tasks. Crucially, adaptation requires only lightweight task-specific heads, underscoring its flexibility and efficiency.
This work addresses the challenge of dynamically varying sample complexity in multimodal learning, which existing curriculum learning approaches struggle to accommodate as models evolve. It introduces, for the first time, Partial Information Decomposition (PID) theory into multimodal curriculum learning, proposing a progressive curriculum framework that decomposes multimodal interactions into redundant, unique, and synergistic information components. This decomposition enables a dynamic characterization of sample complexity, which in turn drives an adaptive scheduling of training samples aligned with the model’s learning progress. The resulting approach yields an interpretable, training-aware sample selection mechanism that significantly outperforms conventional training strategies and state-of-the-art baselines across multiple multimodal benchmarks, demonstrating both its effectiveness and generalizability.
Existing unified vision-language models struggle to generate interleaved multimodal content, limiting their applicability in tasks such as visual storytelling and step-by-step reasoning. This work proposes a reinforcement learning–based post-training approach that endows models with high-quality multimodal interleaved generation capabilities without requiring large-scale interleaved data. The key innovation lies in the first extension of Group Relative Policy Optimization (GRPO) to multimodal settings, complemented by a hybrid reward mechanism that integrates text relevance, image-text alignment, and structural fidelity, along with process-level rewards to jointly optimize text and image generation. Evaluated on the MMIE and InterleavedBench benchmarks, the method significantly improves generation quality, coherence, and structural faithfulness, demonstrating its effectiveness and generalization capability.
This work proposes ERNIE 5.0, the first trillion-parameter native autoregressive multimodal foundation model capable of unified processing of text, images, video, and audio. To address the challenge of efficient deployment under resource constraints, the model employs an ultra-sparse mixture-of-experts (MoE) architecture with a modality-agnostic expert routing mechanism and is trained from scratch using a unified “next group of tokens” prediction objective. A novel elastic training paradigm is introduced, enabling the simultaneous learning of a family of prunable submodels within a single pretraining run, with dynamic adjustment of depth, expert capacity, and sparsity. This approach systematically resolves the stability and efficiency challenges of multimodal reinforcement learning under ultra-sparse MoE settings, achieving balanced and state-of-the-art performance across both multimodal understanding and generation tasks.
This work addresses the challenges of task interference and high memory overhead from storing task-specific prompts in continual video question answering (VideoQA). To this end, the authors propose HyperTokens, a Transformer-based dynamic token generator that produces fine-tuning tokens on demand, enabling explicit control over prompt updates under a fixed memory budget. By incorporating a forward-looking regularization term to suppress sharp, task-specific optimization directions and combining causal modeling with a mutual information proxy loss, the method encourages convergence toward flat minima across modalities and promotes robust continual transfer. Evaluated on two standard continual VideoQA benchmarks, HyperTokens significantly improves average accuracy while reducing forgetting. It also demonstrates strong robustness under a newly introduced cross-modal continual transfer protocol from ImageQA to VideoQA.
This work systematically investigates the core mechanisms of modality interaction in multimodal pretraining, focusing on knowledge transfer, synergistic effects, and fusion timing. Through experiments on both synthetic and large-scale real-world datasets, it provides the first empirical evidence of asymmetric cross-modal knowledge flow and demonstrates that data complexity governs whether modalities exhibit synergy or competition. The study further validates that early unified fusion consistently outperforms late alignment. Leveraging an architecture featuring shared attention and normalization layers with modality-specific feedforward components, the proposed approach is evaluated on a 13.5B mixture-of-experts model trained on 2 trillion tokens, confirming its effectiveness. Additionally, the paper introduces a highly efficient pretraining strategy that achieves strong generative performance using only 5% of the typical computational budget.