MixMAS: A Framework for Sampling-Based Mixer Architecture Search for Multimodal Fusion and Learning

📅 2024-12-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Multimodal fusion faces challenges in adapting heterogeneous data (e.g., images, text, audio) and incurs high costs in manually designing effective architectures. Method: This paper proposes SAMAS, a sampling-driven multimodal mixer architecture search framework. SAMAS introduces the first end-to-end joint optimization paradigm for multimodal learning, simultaneously searching for optimal MLP-based mixer structures, modality-specific encoder combinations, and fusion functions. It employs a lightweight micro-benchmark–based sampling evaluation mechanism to accelerate architecture assessment. Contribution/Results: SAMAS achieves 3–5× higher search efficiency than reinforcement learning or evolutionary algorithms. It attains state-of-the-art fusion performance across multiple standard multimodal benchmarks—including MM-IMDB, CMU-MOSEI, and UR-FUNNY—while substantially reducing reliance on manual architectural design and lowering computational overhead.

Technology Category

Machine Learning: Multimodal LearningComputer Vision: Multi-modal VisionSearch and Optimization: Sampling/Simulation-based Search

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Choosing a suitable deep learning architecture for multimodal data fusion is a challenging task, as it requires the effective integration and processing of diverse data types, each with distinct structures and characteristics. In this paper, we introduce MixMAS, a novel framework for sampling-based mixer architecture search tailored to multimodal learning. Our approach automatically selects the optimal MLP-based architecture for a given multimodal machine learning (MML) task. Specifically, MixMAS utilizes a sampling-based micro-benchmarking strategy to explore various combinations of modality-specific encoders, fusion functions, and fusion networks, systematically identifying the architecture that best meets the task's performance metrics.
Problem

Research questions and friction points this paper is trying to address.

Multi-modal Learning
Deep Learning Methodology
Optimal Network Structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

MixMAS
Auto-Exploration
Multi-Modal Data
🔎 Similar Papers
No similar papers found.
Aix-Marseille Univ | IRD
A
Abdelmadjid Chergui
Higher School of Computer Science, 8 Mai 1945, SBA, Algeria
G
Grigor Bezirganyan
Aix-Marseille Univ, LIS, CNRS, Marseille, France
S
Sana Sellami
Aix-Marseille Univ, LIS, CNRS, Marseille, France
L
Laure Berti-Équille
IRD, ESPACE-DEV, Montpellier, France
Sébastien Fournier
Sébastien Fournier
Aix-Marseille Univ, LIS, CNRS, Marseille, France