adaptive modality gating

Designs and implements modules that compute and apply adaptive, per-modality or pairwise gating weights for multimodal feature fusion, using confidence estimates, feature interactions, or learned policies to dynamically suppress noisy or unreliable modality contributions. Builds and evaluates fusion architectures and analyses that produce sample- or pair-specific fusion weightings to stabilize cross-modal integration and improve robustness under varying input quality.

adaptivemodalitygating

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.45
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Beyond Simple Fusion: Adaptive Gated Fusion for Robust Multimodal Sentiment Analysis

Oct 02, 2025
HW
Han Wu
🏛️ University of Macau | iFLYTEK Co., Ltd.

Multimodal sentiment analysis (MSA) suffers from poor fusion robustness and weak fine-grained sentiment discrimination due to heterogeneous modality quality, noise corruption, and semantic conflicts across modalities. To address these challenges, this paper proposes the Adaptive Gating Fusion Network (AGFN), which introduces a novel dual-gating mechanism grounded in information entropy and modality importance. AGFN dynamically weights multimodal features during both unimodal encoding and cross-modal interaction, explicitly suppressing interference from low-quality modalities. Furthermore, feature distribution visualization guides low-redundancy, high-robustness multimodal representation learning. Extensive experiments on CMU-MOSI and CMU-MOSEI demonstrate that AGFN consistently outperforms state-of-the-art baselines, achieving significant improvements in accuracy—particularly in detecting subtle sentiment shifts—and enhancing model generalizability.

Addressing modality quality variations in multimodal sentiment analysisEnhancing robustness in discerning subtle emotional nuancesMitigating noisy and conflicting modality impacts on predictions

This work addresses critical limitations in existing dynamic multimodal fusion approaches, which rely on heuristic metrics to assess modality quality, fail under extreme noise conditions, and overlook inherent inter-modality dependency biases—leading to doubly suppressed learning of challenging modalities. To overcome these issues, the paper proposes an Unbiased Dynamic Multimodal Learning (UDML) framework that innovatively integrates controlled noise injection with uncertainty prediction to construct a noise-aware uncertainty estimator. Furthermore, UDML explicitly quantifies dependency bias through a modality dropout strategy, enabling adaptive and unbiased weighting of modality contributions. Extensive experiments across multiple multimodal benchmark tasks demonstrate that UDML significantly outperforms both static and state-of-the-art dynamic fusion methods, exhibiting strong effectiveness, broad applicability, and robust generalization capability.

dynamic multimodal fusionmodality biasmodality quality assessment

Multimodal mixture-of-experts (MoE) models suffer from performance degradation when incorporating noise-dominant auxiliary modalities; existing approaches—such as partial information decomposition—struggle to scale beyond two modalities and lack instance-level dynamic modulation. This paper proposes a parameter-free, two-tier weight adaptation framework: an upper tier estimates modality-level importance via inter-modal mutual information, while a lower tier enables fine-grained, instance-specific fusion through KL-divergence-based weighting. To our knowledge, this is the first method enabling non-parametric, dynamic fusion across arbitrarily many modalities without architectural modifications or additional parameters. Evaluated on sentiment regression and clinical classification tasks, it achieves significant improvements—12.3% reduction in mean absolute error (MAE) for regression and 3.8–5.1% gains in multiclass accuracy—demonstrating strong scalability, robustness to modality noise, and cross-task generalization.

Dynamically adjusting modality importance during trainingHandling noise from additional modalities beyond twoStabilizing variance in multimodal model integration

MoPE: Mixture of Prompt Experts for Parameter-Efficient and Scalable Multimodal Fusion

Mar 14, 2024
RJ
Ruixiang Jiang
🏛️ The Hong Kong Polytechnic University | Peng Cheng Lab

Weak adaptability and limited expressiveness of prompt-based multimodal fusion methods lead to suboptimal performance. To address this, we propose the Dynamic Expert Prompt (DEP) framework: it decomposes static prompts into expert prompt modules that are dynamically routed per instance based on instance-specific features and modality-pair priors—achieving, for the first time, instance-level prompt decomposition and adaptive selection. We further introduce a routing regularization mechanism to encourage expert specialization, thereby enhancing interpretability and generalization. With only 0.8% trainable parameters, DEP achieves state-of-the-art performance on six cross-modal datasets spanning four modalities, matching full fine-tuning accuracy while improving parameter efficiency by over 120×.

Adaptability and ExpressivenessParameter EfficiencyPrompt-based Multimodal Fusion

Interpretation on Multi-modal Visual Fusion

Aug 19, 2023
HC
Hao Chen
🏛️ Southeast University | Beijing University of Technology

RGB-D multimodal fusion mechanisms have long suffered from poor interpretability, and the fundamental nature of cross-modal complementarity remains unclear. Method: This paper establishes the first interpretability analysis framework specifically targeting the fusion process, introducing a joint metric of semantic variance and feature similarity to systematically characterize cross-modal representation consistency, intra-modal evolutionary patterns, and collaborative optimization logic. Through cross-layer feature comparison and quantitative semantic analysis, we identify a prevalent imbalance between consistency and specificity in mainstream fusion strategies. Contribution/Results: We formalize a “specificity-driven inference under consistency constraints” principle that explicates cross-modal complementarity. Our framework provides both theoretical foundations and a verifiable evaluation paradigm for designing trustworthy, generalizable multimodal fusion models.

Analyzing feature consistency and specialty across RGB-D modalitiesDeveloping improved fusion strategies for multi-modal RGB-D learningUnderstanding complementary and fusion mechanisms in RGB-D models

Latest Papers

What's happening recently
View more

This work addresses the challenge that certain modalities may introduce interference under specific inputs in multimodal fusion, thereby degrading model performance. To mitigate this issue, the authors propose a plug-and-play pre-fusion calibration module that leverages cross-modal summary contrast to extract supportive and conflicting cues, generating instance-level and dimension-level modulation signals. These signals dynamically enhance beneficial features while suppressing misleading information. The method achieves, for the first time, fine-grained, conflict-aware modulation prior to fusion and is compatible with both sequential and convolutional architectures. It consistently improves performance across five benchmark tasks—including emotion understanding and action recognition—and demonstrates enhanced robustness and consistency under modality missingness and data perturbations.

contextual adjustmentcross-modal conflictmodality calibration

This work addresses the limitations of early fusion—lacking modularity—and late fusion—neglecting cross-modal interactions—in multimodal sentiment recognition by proposing xgaf, an adaptive fusion method grounded in TreeSHAP attribution. xgaf employs a tree-based mixture-of-experts architecture to dynamically weight unimodal and cross-modal experts. Through systematic evaluation of various SHAP reduction strategies, the study identifies sum-abs as particularly effective, as it preserves total attribution magnitude while enhancing performance. The primary performance gain stems from incorporating trimodal experts rather than from complex routing mechanisms. On the MELD and CMU-MOSEI datasets, xgaf achieves weighted F1 scores of 0.5983 and 0.6519, respectively—significantly outperforming late fusion and matching or slightly surpassing early fusion—while simultaneously maintaining modularity and effectively modeling cross-modal interactions.

cross-modal interactionemotion recognitionfeature dimensionality

This work addresses the challenge in multimodal sentiment analysis where modality-specific signal refinement and cross-modal interaction modeling often interfere with each other due to conflicting optimization objectives. To resolve this, the authors propose SeRIn, a novel architecture that decouples modality separation and cross-modal interaction into structured priors through a three-stage pipeline—separation, refinement, and integration—processing unimodal representations and cross-modal interactions via independent pathways before fusing them at the prediction stage. Notably, SeRIn adaptively adjusts modality weights without requiring explicit supervision. The method achieves state-of-the-art performance on the CH-SIMS and CMU-MOSEI benchmarks, yielding significant improvements across all evaluation metrics.

cross-modal interactionmodality-specific refinementmultimodal fusion

Existing approaches struggle to effectively model the dynamic interplay among redundant, unique, and synergistic information at the sample level in multimodal learning. This work proposes a novel information-theoretic decomposition–based paradigm for multimodal interaction learning, offering the first systematic analysis of the importance of sample-level interactions. By employing a variational architecture, the method explicitly disentangles these three interaction components and integrates a component-aware fine-tuning strategy to adaptively leverage them. Extensive experiments across diverse tasks and architectures consistently demonstrate the superiority of the proposed approach over current state-of-the-art methods, confirming its effectiveness and generality in finely modeling sample-level multimodal interactions.

dynamic interactioninformation decompositionmultimodal interaction

This work addresses the performance degradation in multimodal classification caused by modality imbalance by proposing deep ensembling as an alternative to explicit modality fusion, achieving effective multimodal classification through the combination of unimodal networks. The key contributions include the first demonstration that superior performance can be attained without explicit fusion, a heuristic strategy for allocating the number of ensemble models based on each modality’s predictive capability, and the construction of a controllable synthetic multimodal data framework with fitted scaling laws. Experiments show that, under identical parameter budgets, the proposed method significantly outperforms state-of-the-art late-fusion and intermediate-fusion approaches on both real-world and synthetic datasets, while the derived scaling laws reveal an asymptotic upper bound on ensemble performance.

deep ensembleslate-fusionmodality fusion

Hot Scholars

SW

Shicai Wei

University of Electronic Science and Technology of China, UESTC
multimodal learning
JT

Jin Tang

Anhui University
Computer visionintelligent video analysis
YM

Yanda Meng

University of Exeter
Medical Image Analysis
HM

Haoran Ma

PhD Student, University of California, Los Angeles
Computer SystemsSoftware Engineering
KA

Kannan Achan

Walmartlabs
machine learningartificial intelligencegenerative modeling