MultiModalPFN: Extending Prior-Data Fitted Networks for Multimodal Tabular Learning

📅 2026-02-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of existing TabPFN models in effectively integrating heterogeneous non-tabular modalities—such as images and text—which hinders their applicability in multimodal domains like healthcare and marketing. To overcome this, we propose a unified multimodal learning framework that employs modality-specific encoders and projectors to map non-tabular data into tabular-compatible tokens. Furthermore, we introduce a multi-head gated MLP and a cross-attention pooling mechanism to mitigate attention imbalance across modalities and enhance contextual information extraction. Experimental results demonstrate that our approach significantly outperforms current state-of-the-art methods on both medical and general-purpose multimodal datasets, confirming its effectiveness and scalability in jointly leveraging tabular and non-tabular features.

Technology Category

Machine Learning: Multimodal LearningComputer Vision: Multi-modal VisionNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphs
📝 Abstract
Recently, TabPFN has gained attention as a foundation model for tabular data. However, it struggles to integrate heterogeneous modalities such as images and text, which are common in domains like healthcare and marketing, thereby limiting its applicability. To address this, we present the Multi-Modal Prior-data Fitted Network (MMPFN), which extends TabPFN to handle tabular and non-tabular modalities in a unified manner. MMPFN comprises per-modality encoders, modality projectors, and pre-trained foundation models. The modality projectors serve as the critical bridge, transforming non-tabular embeddings into tabular-compatible tokens for unified processing. To this end, we introduce a multi-head gated MLP and a cross-attention pooler that extract richer context from non-tabular inputs while mitigates attention imbalance issue in multimodal learning. Extensive experiments on medical and general-purpose multimodal datasets demonstrate that MMPFN consistently outperforms competitive state-of-the-art methods and effectively exploits non-tabular modalities alongside tabular features. These results highlight the promise of extending prior-data fitted networks to the multimodal setting, offering a scalable and effective framework for heterogeneous data learning. The source code is available at https://github.com/too-z/MultiModalPFN.
Problem

Research questions and friction points this paper is trying to address.

multimodal learning
tabular data
heterogeneous modalities
foundation model
data integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

MultiModalPFN
modality projector
cross-attention pooler
tabular-compatible tokens
gated MLP
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Wall Kim
Samsung Electronics, Hwaseong, South Korea
C
Chaeyoung Song
Seoul National University of Science and Technology, Seoul, South Korea
Hanul Kim
Hanul Kim
Seoul National University of Science and Technology
Computer VisionMachine Learning