Score
Design, build, or analyze fusion modules that first estimate reliability or uncertainty scores for modalities, sensors, views, regions, pixels, or tokens and convert those scores into adaptive weights, gates, or confidences. Integrate those weights into attention/cross-attention, coarse-to-fine pyramid, residual, or consensus fusion pipelines to emphasize reliable cues, suppress corrupted or occluded evidence, aggregate complementary contributions, and preserve structural uncertainty across scales while avoiding hard visibility thresholds.
This work addresses the challenge of effectively fusing heterogeneous thermal and visible-spectrum sensors for reliable drone detection, given their disparities in resolution, viewpoint, and field of view. To this end, two multimodal fusion strategies are proposed: RGIF, which leverages ECC registration and guided filtering, and RGMAF, which integrates affine/optical-flow alignment with a reliability-weighted attention mechanism. By incorporating alignment-awareness and reliability gating, the methods adaptively combine the high contrast of thermal imagery with the rich detail of visible data, thereby mitigating poor spatial correspondence and annotation inconsistencies. Evaluated on the MMFW-UAV dataset, RGIF achieves a mAP@50 of 97.65%, while RGMAF attains the highest recall of 98.64%, both significantly outperforming single-modality baselines.
This work addresses the limitation of existing RGB-infrared object detection methods, which discard spectral statistical information during cross-modal fusion, thereby precluding reliable assessment of fusion quality. To overcome this, the study introduces a novel framework that explicitly extracts a parameter-free 7-dimensional spectral reliability descriptor and reuses it throughout subsequent computations. Specifically, it proposes Spectral Reliability Fusion (SRF) and Reliability-Conditioned Expert Routing (RCER) mechanisms to enable adaptive gated fusion and sparse expert selection. Evaluated on the DroneVehicle dataset under six synthetic degradation scenarios, the method achieves an average retention rate of 95.0% and improves mAP50 by 5.2 and 5.3 points in daytime and nighttime settings, respectively, significantly outperforming content-only baseline approaches.
This work addresses the unclear role of reliability scores in existing quality-aware multimodal fusion methods—specifically, whether these scores genuinely guide model decisions. To investigate this, the authors propose a leakage-safe diagnostic approach: during inference, the model is frozen and reliability scores are shuffled across test samples to assess their actual impact on final predictions. This method effectively distinguishes whether the scores actively drive fusion decisions or merely correlate with performance. Experiments on the StressID and CMU-MOSEI datasets reveal that shuffling scores has negligible effect on performance in real-world scenarios; significant gains from the fusion mechanism occur only when the reliability scores accurately predict the correctness of individual modalities.
This work addresses the challenge of drone detection in misaligned RGB and infrared remote sensing imagery, where spatial misregistration, small target sizes, and complex backgrounds degrade performance. To tackle this, the authors propose LER-YOLO, a novel framework featuring an uncertainty-aware object alignment module that generates spatial reliability maps. These maps guide a sparse Mixture-of-Experts (MoE) fusion module to adaptively select among RGB-dominant, infrared-dominant, or interactive fusion pathways, enabling reliability-driven cross-modal feature integration. Without increasing model capacity, the method achieves 89.7±0.2% AP50 on the MBU benchmark, with a peak performance of 89.9%, substantially outperforming existing approaches in misaligned multimodal scenarios and demonstrating the efficacy of reliability-guided routing for cross-modal fusion.
To address the severe degradation in 3D detection performance caused by single-modality sensor failure (LiDAR or camera) in autonomous driving, this paper proposes ReliFusion—a reliability-driven fusion framework operating in the bird’s-eye-view (BEV) space. Its core contribution is the first explicit quantification of per-modality reliability, integrated with spatio-temporal feature aggregation (STFA) and confidence-weighted mutual cross-attention (CW-MCA) to enable fault-aware, dynamic cross-modal fusion. By adaptively modulating modality weights at the BEV feature level, ReliFusion significantly enhances system robustness under sensor degradation. Evaluated on the nuScenes benchmark, ReliFusion surpasses all state-of-the-art methods, particularly maintaining high accuracy under LiDAR field-of-view occlusion and severe sensor faults. This demonstrates the critical role of explicit reliability modeling in multi-modal fusion for safety-critical perception systems.
Existing multimodal fusion models often lack robustness under noisy or uninformative data and fail to dynamically assess data quality or produce reliable confidence estimates, limiting their applicability in high-stakes clinical settings. To address these challenges, this work proposes the Adaptive Confidence-weighted Extension (ACE) framework, which uniquely integrates intra-modality correlation–driven complementary modality generation with a dual-level dynamic confidence mechanism. This enables adaptive weighting of modality reliability and outputs a global trust score. Evaluated on four multi-omics datasets—BRCA, KIPAN, LGG, and ROSMAP—ACE significantly outperforms current methods, achieving notable improvements in both classification accuracy and confidence calibration, thereby enhancing model robustness and clinical trustworthiness.
This work addresses the problem of miscalibrated confidence in multimodal fusion caused by missing modalities. It proposes Modal-Conditional Conformal Fusion (MCCF), a method that integrates evidential deep learning with Dempster–Shafer theory. During training, MCCF simulates missing modalities via random modality dropout, allowing absent modalities to contribute vacuous evidence automatically. By incorporating Mondrian conformal prediction, MCCF provides finite-sample coverage guarantees for any non-empty subset of available modalities at test time—without requiring imputation. To the best of our knowledge, MCCF is the first approach to achieve formally calibrated uncertainty under arbitrary modality availability, while also decomposing evidence to yield modality-level nullity scores for uncertainty attribution. Experiments on synthetic data and three real-world benchmarks demonstrate that MCCF consistently attains target coverage, substantially narrows the coverage gap between full and partial modalities, and preserves predictive accuracy.
This work addresses the hard-class reliability problem (HCRP) in object detection, where long-tailed minority classes consistently fail in critical scenarios. To tackle this, the authors propose ED-CCF, a decision-level inference framework that reformulates output fusion into an auditable and statistically guaranteed paradigm. ED-CCF introduces a four-state error classification scheme and a class-conditional dynamic calibration mechanism, which activates a calibration pathway only when sufficient empirical evidence is available to precisely correct hard-class errors. By integrating Bonferroni-corrected Wilcoxon significance tests with a Pareto-optimality preservation strategy, ED-CCF achieves a 22.4% improvement in mAP50 for the critical vulnerable class cz (from 0.089 to 0.109) on a 600-image benchmark, while slightly increasing the overall mAP50 to 0.585. Across 50 subset trials, it attains a 96% win rate (p<0.05), significantly enhancing robustness without compromising performance on dominant classes.
This work addresses the limitations of early fusion—lacking modularity—and late fusion—neglecting cross-modal interactions—in multimodal sentiment recognition by proposing xgaf, an adaptive fusion method grounded in TreeSHAP attribution. xgaf employs a tree-based mixture-of-experts architecture to dynamically weight unimodal and cross-modal experts. Through systematic evaluation of various SHAP reduction strategies, the study identifies sum-abs as particularly effective, as it preserves total attribution magnitude while enhancing performance. The primary performance gain stems from incorporating trimodal experts rather than from complex routing mechanisms. On the MELD and CMU-MOSEI datasets, xgaf achieves weighted F1 scores of 0.5983 and 0.6519, respectively—significantly outperforming late fusion and matching or slightly surpassing early fusion—while simultaneously maintaining modularity and effectively modeling cross-modal interactions.
This work addresses the challenge of balancing energy efficiency, latency, and reliability in multimodal edge intelligence systems, where existing approaches either rely on cloud-based fusion or employ single-modality near-sensor filtering that neglects cross-modal dependencies, often resulting in redundant data transmission or missed events. To overcome these limitations, the authors propose FusionSense, a novel framework featuring a three-stage near-sensor learning mechanism. It leverages server-side multimodal models to generate “Filter-out-safe” (FoS) labels that quantify modality necessity, guiding lightweight edge classifiers to jointly optimize computation and communication overhead. Furthermore, it introduces a linearly scalable edge fusion decision mechanism. Evaluated on RGB+depth/LiDAR scenarios, FusionSense achieves up to 33× lower energy consumption at 1% FoS, reduces quality loss by 92.3% under 30% compression, and improves energy-efficiency gains by approximately 1.5× compared to baseline methods.