quality-aware moe fusion

Designs and implements multimodal fusion mechanisms that estimate per-modality quality or uncertainty (including heteroscedastic, input-dependent noise) and use those estimates to weight, gate, or modulate each modality's contribution. This involves modeling modality-specific error distributions, learning quality-aware weighting or attention functions, and producing fused latent/contextual embeddings or calibrated downstream predictions.

quality-awaremoefusion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses critical limitations in existing dynamic multimodal fusion approaches, which rely on heuristic metrics to assess modality quality, fail under extreme noise conditions, and overlook inherent inter-modality dependency biases—leading to doubly suppressed learning of challenging modalities. To overcome these issues, the paper proposes an Unbiased Dynamic Multimodal Learning (UDML) framework that innovatively integrates controlled noise injection with uncertainty prediction to construct a noise-aware uncertainty estimator. Furthermore, UDML explicitly quantifies dependency bias through a modality dropout strategy, enabling adaptive and unbiased weighting of modality contributions. Extensive experiments across multiple multimodal benchmark tasks demonstrate that UDML significantly outperforms both static and state-of-the-art dynamic fusion methods, exhibiting strong effectiveness, broad applicability, and robust generalization capability.

dynamic multimodal fusionmodality biasmodality quality assessment

This work addresses the challenge that certain modalities may introduce interference under specific inputs in multimodal fusion, thereby degrading model performance. To mitigate this issue, the authors propose a plug-and-play pre-fusion calibration module that leverages cross-modal summary contrast to extract supportive and conflicting cues, generating instance-level and dimension-level modulation signals. These signals dynamically enhance beneficial features while suppressing misleading information. The method achieves, for the first time, fine-grained, conflict-aware modulation prior to fusion and is compatible with both sequential and convolutional architectures. It consistently improves performance across five benchmark tasks—including emotion understanding and action recognition—and demonstrates enhanced robustness and consistency under modality missingness and data perturbations.

contextual adjustmentcross-modal conflictmodality calibration

This work addresses the unclear role of reliability scores in existing quality-aware multimodal fusion methods—specifically, whether these scores genuinely guide model decisions. To investigate this, the authors propose a leakage-safe diagnostic approach: during inference, the model is frozen and reliability scores are shuffled across test samples to assess their actual impact on final predictions. This method effectively distinguishes whether the scores actively drive fusion decisions or merely correlate with performance. Experiments on the StressID and CMU-MOSEI datasets reveal that shuffling scores has negligible effect on performance in real-world scenarios; significant gains from the fusion mechanism occur only when the reliability scores accurately predict the correctness of individual modalities.

decision-level dependencemodality weightingmultimodal fusion

This work addresses the performance degradation in multimodal classification caused by modality imbalance by proposing deep ensembling as an alternative to explicit modality fusion, achieving effective multimodal classification through the combination of unimodal networks. The key contributions include the first demonstration that superior performance can be attained without explicit fusion, a heuristic strategy for allocating the number of ensemble models based on each modality’s predictive capability, and the construction of a controllable synthetic multimodal data framework with fitted scaling laws. Experiments show that, under identical parameter budgets, the proposed method significantly outperforms state-of-the-art late-fusion and intermediate-fusion approaches on both real-world and synthetic datasets, while the derived scaling laws reveal an asymptotic upper bound on ensemble performance.

deep ensembleslate-fusionmodality fusion

Multimodal Alignment and Fusion: A Survey

Nov 26, 2024
SL
Songtao Li
🏛️ Northeastern University | Peking University

This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.

Addressing cross-modal misalignment and computational bottlenecks challengesExploring applications from social media to medical imagingSurveying multimodal alignment and fusion techniques in machine learning

Latest Papers

What's happening recently
View more

This work addresses the common yet often separately treated challenges of missing modalities and noisy data in real-world multimodal scenarios. To unify these issues under a single framework, the authors propose Unified Modality Quality (UMQ), which models both missingness and noise as manifestations of low-quality modalities. UMQ introduces a ranking-guided quality estimation strategy, modality baseline representations, sample-specific cross-modal fusion, and a quality-aware mixture-of-experts routing mechanism to jointly optimize learning from degraded modalities. Extensive experiments on multiple multimodal sentiment analysis benchmarks demonstrate that UMQ consistently outperforms state-of-the-art methods across various settings—including complete, missing, and noisy modalities—highlighting its robustness and effectiveness in handling real-world multimodal imperfections.

low-quality multimodal datamissing modalitiesmodality quality

Existing performance-degradation-based methods for modality contribution assessment struggle to disentangle a modality’s unique information content from its synergistic interaction effects—particularly under cross-attention architectures, where they face fundamental limitations. To address this, we propose the first representation-level quantification framework grounded in Partial Information Decomposition (PID), which rigorously decomposes multimodal representations into unique, redundant, and synergistic information components. Our method integrates PID theory with the Iterative Proportional Fitting Procedure (IPFP), enabling layer-wise and cross-dataset contribution inference without retraining. Experiments demonstrate that our framework substantially enhances both the interpretability and accuracy of contribution analysis. It provides a principled, fine-grained separation of each modality’s independent and interactive contributions, establishing a new paradigm for multimodal model diagnosis and architecture design.

Developing scalable inference-only analysis without model retraining requirementsDistinguishing inherent information from synergistic interactions between modalitiesQuantifying modality contributions in multimodal models using disentangled representations

This work addresses the challenge in multimodal sentiment analysis that different modalities exhibit significantly varying reliability at the utterance level due to factors such as occlusion, noise, or transcription errors, which can adversely affect traditional fusion approaches. To mitigate this issue, the authors propose MRUF, a method that enables adaptive fusion through multi-granularity routing and uncertainty-aware calibration. Specifically, MRUF incorporates modality importance supervision derived from leave-one-out error increases, an inverse-variance reweighting mechanism for uncertainty-based gating, and a modality-invariant contrastive alignment strategy. Experimental results demonstrate that MRUF consistently outperforms strong baselines on both aligned and unaligned settings of the CMU-MOSI and CMU-MOSEI datasets, effectively down-weighting contributions from high-uncertainty modalities and thereby enhancing model robustness.

fusionmodality qualitymodality reliability

This work addresses the challenge in multimodal sentiment analysis where modality-specific signal refinement and cross-modal interaction modeling often interfere with each other due to conflicting optimization objectives. To resolve this, the authors propose SeRIn, a novel architecture that decouples modality separation and cross-modal interaction into structured priors through a three-stage pipeline—separation, refinement, and integration—processing unimodal representations and cross-modal interactions via independent pathways before fusing them at the prediction stage. Notably, SeRIn adaptively adjusts modality weights without requiring explicit supervision. The method achieves state-of-the-art performance on the CH-SIMS and CMU-MOSEI benchmarks, yielding significant improvements across all evaluation metrics.

cross-modal interactionmodality-specific refinementmultimodal fusion

This work addresses the practical challenges in multimodal learning caused by missing or redundant modalities and the lack of theoretical understanding of how modality selection affects performance. It presents the first joint analysis of how the number of modalities and feature granularity influence generalization. By constructing a hierarchical structure of function classes corresponding to different modality subsets and leveraging pairwise complexity measures, the study derives generalization error bounds that quantify the discrepancy between the learned mapping and the true underlying mapping. The theoretical results demonstrate that fine-grained modality features effectively reduce hypothesis space complexity and enhance modality complementarity, thereby improving both convergence rates and prediction accuracy. These findings provide a rigorous theoretical foundation for designing effective multimodal learning systems.

generalization guaranteesmetric learningmodality selection

Hot Scholars

YO

Yassine Ouali

Samsung AI Cambridge
Machine LearningDeep Learning
AB

Adrian Bulat

Samsung AI Cambridge
Computer VisionDeep LearningMachine LearningArtificial Intelligence
LC

Lequn Chen

University of Washington
Distributed SystemsMachine Learning SystemsOperating Systems
SK

Seung Ki Moon

Associate Professor of Mechanical and Aerospace Engineering, Nanyang Technological University
Product family and platform designadditive manufacturing and 3D printingoptimizationdigital twins
JK

Jinman Kim

School of Computer Science, University of Sydney
Medical VisualisationMedical Image SegmentationTelehealth