composable multimodal pretraining

Designs and implements training frameworks, model architectures, and optimization procedures that pretrain models over multiple modalities in a unified, modular, and composable way—using interleaved autoregressive schedules, native multimodal objectives, and unified training strategies—to jointly learn representations while preserving core pretrained knowledge. Builds mechanisms and interfaces for non‑destructive addition of modalities and plug‑and‑play expert composition that support both generative and discriminative objectives and enable modular expansion without catastrophic interference.

composablemultimodalpretraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.55
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work systematically investigates the core mechanisms of modality interaction in multimodal pretraining, focusing on knowledge transfer, synergistic effects, and fusion timing. Through experiments on both synthetic and large-scale real-world datasets, it provides the first empirical evidence of asymmetric cross-modal knowledge flow and demonstrates that data complexity governs whether modalities exhibit synergy or competition. The study further validates that early unified fusion consistently outperforms late alignment. Leveraging an architecture featuring shared attention and normalization layers with modality-specific feedforward components, the proposed approach is evaluated on a 13.5B mixture-of-experts model trained on 2 trillion tokens, confirming its effectiveness. Additionally, the paper introduces a highly efficient pretraining strategy that achieves strong generative performance using only 5% of the typical computational budget.

early unificationknowledge flowmodality interaction

This work addresses the challenge of catastrophic forgetting in foundational models when continually integrating new modalities, which hinders their ability to jointly support generation and comprehension tasks. To overcome this, the authors propose a composable, natively multimodal pretraining framework that maintains general knowledge through globally shared experts while employing plug-in experts for modality-specific capabilities. Central to this approach is the novel Momentum-Anchored Orthogonal Projection (MAOP) mechanism, which leverages optimizer momentum states as semantic anchors to selectively neutralize conflicting gradients during modality expansion, enabling lossless knowledge fusion. The method effectively mitigates forgetting, robustly preserves language and visual understanding performance, enhances image generation quality, and elicits beneficial cross-modal synergies.

catastrophic forgettingfoundation modelsgradient conflict

Improving DNN Modularization via Activation-Driven Training

Nov 01, 2024
TN
Tuan Ngo
🏛️ University of Southern California

To address the high retraining cost and accumulating technical debt when adapting deep neural networks (DNNs) to new tasks, this paper proposes MODA, an activation-driven modular training framework. Unlike existing modularization methods relying on masks or post-hoc processing, MODA introduces an end-to-end differentiable optimization over activation space, jointly pursuing intra-class aggregation, inter-class separation, and module compactness—enabling natural, non-overlapping modular decomposition across all layers (not limited to convolutional layers). It incorporates activation regularization and class-aware distribution constraints, allowing plug-and-play module replacement without fine-tuning. Experiments demonstrate that MODA reduces training time by 29%, decreases module parameters by 58%, lowers weight overlap by 71%, and incurs zero accuracy loss. After module replacement, target-class accuracy improves by 12%, while accuracy on other classes varies by less than 0.5%.

Improving modularization efficiency without auxiliary masksMinimizing weight overlaps and accuracy losses in modulesReducing technical debt and retraining costs in DNNs

Harmony: A Unified Framework for Modality Incremental Learning

Apr 17, 2025
YS
Yaguang Song
🏛️ Peng Cheng Laboratory | Institute of Automation, Chinese Academy of Sciences | University of Chinese Academy of Sciences

This work addresses the challenge of incremental learning under realistic scenarios where modalities continuously emerge and cross-modal paired data are unavailable, proposing a novel paradigm termed Modality-Incremental Learning (MIL). To enable cross-modal knowledge transfer and long-term retention under stage-wise unimodal training, we design an Adaptive Compatible Feature Modulation mechanism and a Cumulative Modality Bridging mechanism—achieving modality alignment and historical knowledge consolidation without paired supervision. Our unified framework jointly integrates historical modality feature reuse, progressive accumulation of modality-specific knowledge, and dynamic modulation of the shared feature space. Evaluated on the first dedicated MIL benchmark, our method significantly outperforms existing incremental learning approaches, demonstrating strong efficacy, generalizability, and scalability for continual learning over heterogeneous modality sequences.

Address modal discrepancy and knowledge retention in distinct modalitiesDevelop a unified model for incremental learning across evolving modalitiesEnable cross-modal task completion within a single framework

This work addresses the catastrophic forgetting in language understanding tasks that often arises when large multimodal language models are endowed with image generation capabilities, primarily due to gradient conflicts between generative and discriminative objectives. To mitigate this issue, the authors propose a native multimodal mixture-of-experts (MoE) architecture that jointly optimizes generation and understanding within a unified pretraining framework. The approach leverages modality-aware expert decoupling, shared experts as cross-modal semantic bridges, differential learning rates, and early-stage gradient masking—all without introducing any additional parameters. Experimental results demonstrate that the proposed method significantly enhances performance on language understanding benchmarks such as MMLU and OCRBench while simultaneously accelerating convergence in image generation tasks.

catastrophic forgettinggradient conflictimage generation

Latest Papers

What's happening recently
View more

This work addresses the challenge of balancing stability and plasticity in continual learning under sequential data scenarios. Inspired by the modular organization of the human brain, the authors propose MoRe, a novel framework that constructs a theoretically identifiable hierarchical modular structure in representation space. MoRe decomposes knowledge into shared foundational modules and task-specific modules, enabling module reuse, alignment, and expansion. By leveraging temporal delayed dependencies to uncover intrinsic sequence structures and integrating modular learning with identifiability constraints, MoRe achieves structured knowledge organization and protection without requiring explicit task boundaries. Experiments on synthetic benchmarks and activation data from large language models demonstrate that MoRe learns interpretable hierarchical representations and significantly improves the stability-plasticity trade-off in continual learning.

continual learningmodularityplasticity-stability trade-off

This study addresses the limitation of existing multimodal models in adapting to arbitrary modality combinations and prediction tasks by proposing a universal multimodal foundation model. The proposed method introduces a structural multimodal causal model to generate large-scale synthetic data, enabling the model to learn transferable cross-modal correlation patterns rather than relying on modality-specific representations. Furthermore, it incorporates in-context exemplar reasoning to achieve zero-shot generalization and fusion. Experimental evaluations across 18 datasets demonstrate that the proposed model attains competitive performance comparable to task-specific models without requiring task-specific fine-tuning. This work establishes a novel paradigm for constructing modality-agnostic, general-purpose multimodal systems.

arbitrary modalitiesfoundation modelgeneralization

This work addresses the disconnect between multitask pretraining and subsequent continual learning in multimodal models by proposing a scalable sparse Mixture-of-Experts (MoE) framework. The approach introduces a modality-aware router to handle heterogeneous inputs and compresses expert knowledge into a low-rank memory subspace. By expanding only lightweight routers while keeping the backbone capacity fixed, the framework enables efficient task-incremental continual learning. It effectively mitigates catastrophic forgetting and substantially improves parameter efficiency. Extensive experiments on multiple medical multimodal benchmarks demonstrate the model’s ability to continuously adapt to new tasks while preserving pretrained performance.

catastrophic forgettingcontinual learningmodality combinations

Hot Scholars

MM

Ming-Ming Cheng

Professor of Computer Science, Nankai University
Computer VisionComputer GraphicsVisual AttentionSaliency
XS

Xing Sun

Tencent Youtu Lab
LLMMLLMAgent
JZ

Jingren Zhou

Alibaba Group, Microsoft
Cloud ComputingLarge Scale Distributed SystemsMachine LearningQuery Processing
ZL

Zuozhu Liu

Assistant Professor, Zhejiang University/University of Illinois Urbana-Champaign
deep learningvision-language modelsmedical AI
HL

Hongfei Lin

DalianUniversity of Technology
natural language processing,sentimental analysistext miningsocial computing