progressive multimodal training

Designs training curricula and optimization procedures that progressively incorporate multiple data modalities—using hierarchical or stage-wise training, modality-aware objectives, curriculum schedules, and KL-based routing supervision—to stabilize learning of unified multimodal token representations and support variable-length multimodal sequences. Builds adaptive loss-weighting and routing schemes to mitigate heterogeneous-information imbalance, improve joint generation and alignment, and stabilize generation quality across training stages.

progressivemultimodaltraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.49
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the unclear trade-offs among diverse capabilities—such as general visual understanding, structured reasoning, and fine-grained OCR—in multimodal instruction tuning under mixed data regimes, particularly the lack of systematic investigation into how data organization influences these trade-offs. Treating data scheduling as a first-order design variable while holding model architecture and optimization settings fixed, the work compares four strategies: direct mixing, curriculum learning, balanced sampling, and reverse curriculum. Results demonstrate that curriculum-based training—sequencing tasks from general comprehension to specialized skills—achieves superior overall performance and structured reasoning while accelerating convergence. Balanced sampling improves OCR accuracy at the cost of capability imbalance, whereas reverse curriculum degrades performance and induces optimization instability. The findings highlight the critical role of training sequence in shaping the capability distribution of multimodal models.

capability trade-offsdata organizationmultimodal instruction tuning

Existing reinforcement learning post-training methods fail to differentiate among multimodal inputs, leading to high policy gradient variance, slow convergence, and poor robustness to missing modalities or distribution shifts. To address this, this work proposes the first task-level benchmark for minimal modality-combination annotation and introduces MAPO, a modality-aware policy optimization framework that integrates hierarchical batching, adaptive weighting, and a curriculum scheduling mechanism based on signal-combination difficulty. Experiments on MAPLE-bench demonstrate that the proposed approach reduces the accuracy gap between unimodal and multimodal settings by 30.24%, accelerates convergence by a factor of 3.18, and maintains stable performance across diverse modality-missing scenarios.

distribution shiftgradient variancemodality-aware training

This work addresses the optimization instability in multimodal large language models during autoregressive training, which arises from gradient heterogeneity across modalities and hinders scaling to large batch sizes. The authors propose ML-FOP-SOAP, a novel framework that introduces the second-order preconditioning method SOAP into multimodal training for the first time. By integrating Fisher orthogonal projection to mitigate inter-modal competition and incorporating multi-level variance correction with hierarchical folding strategies, the approach enables efficient co-optimization of visual generation and textual understanding under low computational overhead. Evaluated on Janus and Emu3, ML-FOP-SOAP significantly enhances training stability and sample efficiency, supporting stable training with batch sizes up to 8192. It achieves a 1.4× improvement in sample efficiency and reduces training time by 1.5× compared to AdamW.

gradient heterogeneitylarge-batch scalingmodality competition

MIO: A Foundation Model on Multimodal Tokens

Sep 26, 2024
ZW
Z. Wang
🏛️ Beihang University | The Hong Kong Polytechnic University | AIWaves | University of Alberta | University of Manchester | Institute of Automation, Chinese Academy of Sciences | Peking University | University of Waterloo | The Hong Kong University of Science and Technology

Existing large models struggle to jointly model and freely generate interleaved multimodal sequences—such as speech, text, images, and video—within a unified framework. This paper introduces MIO, an end-to-end autoregressive multimodal foundation model, and proposes the first four-stage causal multimodal training paradigm: alignment pretraining → interleaving pretraining → speech-augmented pretraining → multitask supervised fine-tuning. MIO employs discrete multimodal tokenization and unified causal modeling to enable true any-to-any cross-modal understanding and generation. It supports arbitrary input/output modality combinations and long-range interleaved sequence generation, unlocking novel capabilities including chain-of-visual-reasoning and instruction-driven image editing. On comprehensive cross-modal benchmarks, MIO matches or surpasses state-of-the-art dual-modal, any-to-any, and unimodal specialized models. Notably, it achieves significant performance gains on complex interleaved tasks—such as video-text generation and visual-guided instruction generation.

Cross-modal Sequence GenerationLarge Language ModelsMultimodal Information Processing

Everything is a Video: Unifying Modalities through Next-Frame Prediction

Nov 15, 2024
GT
G. Thomas
🏛️ Durham University

Traditional multimodal approaches rely on modality-specific encoders and late fusion, limiting scalability and cross-modal generalization. This work proposes a unified paradigm that reformulates diverse multimodal tasks—including text, image, audio, and video processing—as “next-frame prediction,” with all inputs and outputs represented as serialized video frames, enabling end-to-end, single-model inference. It pioneers the task-redefinition strategy—previously confined to NLP—within multimodal learning, eliminating modality-specific design in favor of a fully modality-agnostic architecture. The framework integrates cross-modal tokenization, sequential frame representation, autoregressive modeling, and a shared Transformer decoder, supporting joint multi-task pretraining. Empirical evaluation demonstrates strong zero-shot and few-shot generalization across text-to-text, image-to-text, video-to-video, video-to-text, and audio-to-text tasks. Crucially, adaptation requires only lightweight task-specific heads, underscoring its flexibility and efficiency.

Enabling seamless knowledge transfer across different tasks and modalitiesOvercoming modality-specific encoder limitations in multimodal learningUnifying diverse modalities via next-frame prediction

Latest Papers

What's happening recently
View more

This work addresses the challenge of dynamically varying sample complexity in multimodal learning, which existing curriculum learning approaches struggle to accommodate as models evolve. It introduces, for the first time, Partial Information Decomposition (PID) theory into multimodal curriculum learning, proposing a progressive curriculum framework that decomposes multimodal interactions into redundant, unique, and synergistic information components. This decomposition enables a dynamic characterization of sample complexity, which in turn drives an adaptive scheduling of training samples aligned with the model’s learning progress. The resulting approach yields an interpretable, training-aware sample selection mechanism that significantly outperforms conventional training strategies and state-of-the-art baselines across multiple multimodal benchmarks, demonstrating both its effectiveness and generalizability.

curriculum learninginformation decompositionmodel evolution

Existing unified vision-language models struggle to generate interleaved multimodal content, limiting their applicability in tasks such as visual storytelling and step-by-step reasoning. This work proposes a reinforcement learning–based post-training approach that endows models with high-quality multimodal interleaved generation capabilities without requiring large-scale interleaved data. The key innovation lies in the first extension of Group Relative Policy Optimization (GRPO) to multimodal settings, complemented by a hybrid reward mechanism that integrates text relevance, image-text alignment, and structural fidelity, along with process-level rewards to jointly optimize text and image generation. Evaluated on the MMIE and InterleavedBench benchmarks, the method significantly improves generation quality, coherence, and structural faithfulness, demonstrating its effectiveness and generalization capability.

multimodal interleaved generationstep-by-step visual reasoningunified multimodal generation

This work proposes ERNIE 5.0, the first trillion-parameter native autoregressive multimodal foundation model capable of unified processing of text, images, video, and audio. To address the challenge of efficient deployment under resource constraints, the model employs an ultra-sparse mixture-of-experts (MoE) architecture with a modality-agnostic expert routing mechanism and is trained from scratch using a unified “next group of tokens” prediction objective. A novel elastic training paradigm is introduced, enabling the simultaneous learning of a family of prunable submodels within a single pretraining run, with dynamic adjustment of depth, expert capacity, and sparsity. This approach systematically resolves the stability and efficiency challenges of multimodal reinforcement learning under ultra-sparse MoE settings, achieving balanced and state-of-the-art performance across both multimodal understanding and generation tasks.

autoregressive modelingelastic trainingmixture-of-experts

This work addresses the challenges of task interference and high memory overhead from storing task-specific prompts in continual video question answering (VideoQA). To this end, the authors propose HyperTokens, a Transformer-based dynamic token generator that produces fine-tuning tokens on demand, enabling explicit control over prompt updates under a fixed memory budget. By incorporating a forward-looking regularization term to suppress sharp, task-specific optimization directions and combining causal modeling with a mutual information proxy loss, the method encourages convergence toward flat minima across modalities and promotes robust continual transfer. Evaluated on two standard continual VideoQA benchmarks, HyperTokens significantly improves average accuracy while reducing forgetting. It also demonstrates strong robustness under a newly introduced cross-modal continual transfer protocol from ImageQA to VideoQA.

catastrophic forgettingContinual VideoQAmultimodal LLMs

This work systematically investigates the core mechanisms of modality interaction in multimodal pretraining, focusing on knowledge transfer, synergistic effects, and fusion timing. Through experiments on both synthetic and large-scale real-world datasets, it provides the first empirical evidence of asymmetric cross-modal knowledge flow and demonstrates that data complexity governs whether modalities exhibit synergy or competition. The study further validates that early unified fusion consistently outperforms late alignment. Leveraging an architecture featuring shared attention and normalization layers with modality-specific feedforward components, the proposed approach is evaluated on a 13.5B mixture-of-experts model trained on 2 trillion tokens, confirming its effectiveness. Additionally, the paper introduces a highly efficient pretraining strategy that achieves strong generative performance using only 5% of the typical computational budget.

early unificationknowledge flowmodality interaction

Hot Scholars

DM

Deyu Meng

Professor, Xi'an Jiaotong University
Machine LearningApplied MathematicsComputer VisionArtificial Intelligence
GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
EC

Ethan Chern

Shanghai Jiao Tong University
Machine LearningNatural Language ProcessingArtificial Intelligence
EB

Egor Bondarev

Associate Professor, Eindhoven University of Technology
computer visionAI3D reconstructionreal-time architectures
SM

Snehashis Majhi

PhD Candidate, STARS Team, INRIA
Computer VisionAbnormal Activity DetectionWeakly-supervised Learning