🤖 AI Summary
This study addresses the limitation of existing multimodal models in adapting to arbitrary modality combinations and prediction tasks by proposing a universal multimodal foundation model. The proposed method introduces a structural multimodal causal model to generate large-scale synthetic data, enabling the model to learn transferable cross-modal correlation patterns rather than relying on modality-specific representations. Furthermore, it incorporates in-context exemplar reasoning to achieve zero-shot generalization and fusion. Experimental evaluations across 18 datasets demonstrate that the proposed model attains competitive performance comparable to task-specific models without requiring task-specific fine-tuning. This work establishes a novel paradigm for constructing modality-agnostic, general-purpose multimodal systems.
📝 Abstract
Making prediction with multimodal data is widely used in diverse scenarios. Existing multimodal fusion models, once deployed, can only handle predefined modalities (e.g., vision, text and audio) and single tasks, making it difficult to quickly adapt to new downstream applications. Therefore, a natural yet rather aggressive question arises, whether there exists a general multimodal fusion model that can be applied to arbitrary modality combinations and arbitrary prediction tasks. We argue that a unified multimodal fusion model should not depend on specific modalities and instead encode transferable patterns of multimodal correlation. To this end, we propose a simple and effective learning paradigm based on training over the generation of large-scale synthetic multimodal datasets with diverse causal structures that formally characterize the generative processes of multimodal data in real world. Building on this framework, we propose the generalized multimodal foundation model, a unified foundation model for generalized multimodal learning. By constructing large-scale synthetic multimodal datasets with diverse correlation patterns, our model encodes transferable multimodal correlations during training and activates appropriate associations through in-context examples during inference. Extensive experiments on 18 real-world datasets spanning 12 modalities and 11 prediction tasks demonstrate that our model achieves competitive performance with specialized models without task-specific adaptation.