Score
Designs and implements training frameworks, model architectures, and optimization procedures that pretrain models over multiple modalities in a unified, modular, and composable way—using interleaved autoregressive schedules, native multimodal objectives, and unified training strategies—to jointly learn representations while preserving core pretrained knowledge. Builds mechanisms and interfaces for non‑destructive addition of modalities and plug‑and‑play expert composition that support both generative and discriminative objectives and enable modular expansion without catastrophic interference.
This paper addresses the challenges in developing general-purpose multimodal large language models (MLLMs) capable of cross-modal coordination across six generative modalities: text, image, music, video, human motion, and 3D objects. To this end, it proposes a novel unified architecture integrating Transformer-based and diffusion-based paradigms, augmented with self-supervised learning (SSL), mixture-of-experts (MoE), reinforcement learning from human feedback (RLHF), and chain-of-thought (CoT) reasoning. The work introduces the first taxonomy covering all six modalities and identifies shared enabling mechanisms for cross-modal transfer. It further argues that structured reasoning and modular decoupling are critical to improving interpretability and generalization. The resulting comprehensive MLLM technology landscape clarifies common bottlenecks and transferable methodologies, providing both theoretical foundations and practical guidelines for building universal, adaptive, and interpretable multimodal systems.
This work systematically investigates the core mechanisms of modality interaction in multimodal pretraining, focusing on knowledge transfer, synergistic effects, and fusion timing. Through experiments on both synthetic and large-scale real-world datasets, it provides the first empirical evidence of asymmetric cross-modal knowledge flow and demonstrates that data complexity governs whether modalities exhibit synergy or competition. The study further validates that early unified fusion consistently outperforms late alignment. Leveraging an architecture featuring shared attention and normalization layers with modality-specific feedforward components, the proposed approach is evaluated on a 13.5B mixture-of-experts model trained on 2 trillion tokens, confirming its effectiveness. Additionally, the paper introduces a highly efficient pretraining strategy that achieves strong generative performance using only 5% of the typical computational budget.
This work addresses the challenge of catastrophic forgetting in foundational models when continually integrating new modalities, which hinders their ability to jointly support generation and comprehension tasks. To overcome this, the authors propose a composable, natively multimodal pretraining framework that maintains general knowledge through globally shared experts while employing plug-in experts for modality-specific capabilities. Central to this approach is the novel Momentum-Anchored Orthogonal Projection (MAOP) mechanism, which leverages optimizer momentum states as semantic anchors to selectively neutralize conflicting gradients during modality expansion, enabling lossless knowledge fusion. The method effectively mitigates forgetting, robustly preserves language and visual understanding performance, enhances image generation quality, and elicits beneficial cross-modal synergies.
To address the high retraining cost and accumulating technical debt when adapting deep neural networks (DNNs) to new tasks, this paper proposes MODA, an activation-driven modular training framework. Unlike existing modularization methods relying on masks or post-hoc processing, MODA introduces an end-to-end differentiable optimization over activation space, jointly pursuing intra-class aggregation, inter-class separation, and module compactness—enabling natural, non-overlapping modular decomposition across all layers (not limited to convolutional layers). It incorporates activation regularization and class-aware distribution constraints, allowing plug-and-play module replacement without fine-tuning. Experiments demonstrate that MODA reduces training time by 29%, decreases module parameters by 58%, lowers weight overlap by 71%, and incurs zero accuracy loss. After module replacement, target-class accuracy improves by 12%, while accuracy on other classes varies by less than 0.5%.
This work addresses the challenge of incremental learning under realistic scenarios where modalities continuously emerge and cross-modal paired data are unavailable, proposing a novel paradigm termed Modality-Incremental Learning (MIL). To enable cross-modal knowledge transfer and long-term retention under stage-wise unimodal training, we design an Adaptive Compatible Feature Modulation mechanism and a Cumulative Modality Bridging mechanism—achieving modality alignment and historical knowledge consolidation without paired supervision. Our unified framework jointly integrates historical modality feature reuse, progressive accumulation of modality-specific knowledge, and dynamic modulation of the shared feature space. Evaluated on the first dedicated MIL benchmark, our method significantly outperforms existing incremental learning approaches, demonstrating strong efficacy, generalizability, and scalability for continual learning over heterogeneous modality sequences.
This work addresses the catastrophic forgetting in language understanding tasks that often arises when large multimodal language models are endowed with image generation capabilities, primarily due to gradient conflicts between generative and discriminative objectives. To mitigate this issue, the authors propose a native multimodal mixture-of-experts (MoE) architecture that jointly optimizes generation and understanding within a unified pretraining framework. The approach leverages modality-aware expert decoupling, shared experts as cross-modal semantic bridges, differential learning rates, and early-stage gradient masking—all without introducing any additional parameters. Experimental results demonstrate that the proposed method significantly enhances performance on language understanding benchmarks such as MMLU and OCRBench while simultaneously accelerating convergence in image generation tasks.
This work addresses the challenge of balancing stability and plasticity in continual learning under sequential data scenarios. Inspired by the modular organization of the human brain, the authors propose MoRe, a novel framework that constructs a theoretically identifiable hierarchical modular structure in representation space. MoRe decomposes knowledge into shared foundational modules and task-specific modules, enabling module reuse, alignment, and expansion. By leveraging temporal delayed dependencies to uncover intrinsic sequence structures and integrating modular learning with identifiability constraints, MoRe achieves structured knowledge organization and protection without requiring explicit task boundaries. Experiments on synthetic benchmarks and activation data from large language models demonstrate that MoRe learns interpretable hierarchical representations and significantly improves the stability-plasticity trade-off in continual learning.
This study addresses the limitation of existing multimodal models in adapting to arbitrary modality combinations and prediction tasks by proposing a universal multimodal foundation model. The proposed method introduces a structural multimodal causal model to generate large-scale synthetic data, enabling the model to learn transferable cross-modal correlation patterns rather than relying on modality-specific representations. Furthermore, it incorporates in-context exemplar reasoning to achieve zero-shot generalization and fusion. Experimental evaluations across 18 datasets demonstrate that the proposed model attains competitive performance comparable to task-specific models without requiring task-specific fine-tuning. This work establishes a novel paradigm for constructing modality-agnostic, general-purpose multimodal systems.
研究统一多模态模型中理解和生成任务的协同效应,通过调整架构和任务知识共享,促进两者从共存到协同。
研究通过模仿果蝇学习记忆系统的层级模块化原理,提出一种轻量级的预训练基础模型适应方法,以解决在线、不确定数据流下的持续学习问题。
This work addresses the disconnect between multitask pretraining and subsequent continual learning in multimodal models by proposing a scalable sparse Mixture-of-Experts (MoE) framework. The approach introduces a modality-aware router to handle heterogeneous inputs and compresses expert knowledge into a low-rank memory subspace. By expanding only lightweight routers while keeping the backbone capacity fixed, the framework enables efficient task-incremental continual learning. It effectively mitigates catastrophic forgetting and substantially improves parameter efficiency. Extensive experiments on multiple medical multimodal benchmarks demonstrate the model’s ability to continuously adapt to new tasks while preserving pretrained performance.