adaptive multi-view fusion

Design and implement adaptive fusion modules that aggregate feature streams from multiple spatial or temporal views (images, frames, 3D feature maps or other per-view embeddings) into a single fused representation. These modules use learned per-view weighting and interaction mechanisms—attention, multi-head/transformer layers, multiplicative gating, latent fusion or data-driven gates—conditioned on viewpoint, distortion or state to suppress noisy inputs, amplify informative streams, and enable cross-view geometric or temporal feature interaction.

adaptivemulti-viewfusion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.71
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$215K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Multimodal image fusion faces two key challenges: gradient conflicts arising from cross-modal parameter sharing, which degrade performance; and modality-specific encoders that improve fusion quality yet harm task generalization. To address these, we propose a unified fusion framework integrating three novel components: semantic-aware channel pruning (to retain discriminative features), geometric affine modulation (to model inter-modal spatial discrepancies), and text-guided channel perturbation (to inject semantic priors and enhance robustness). Our method synergistically leverages pretrained semantic knowledge, channel-level perturbations, and affine transformations—achieving selective feature learning and strong cross-task generalization without introducing modality-specific parameters. Extensive experiments demonstrate consistent and significant improvements over state-of-the-art methods on major fusion benchmarks and downstream detection and segmentation tasks.

Addresses gradient conflicts in unified multi-modality image fusion modelsEnhances feature discriminability while maintaining cross-task generalizationReduces dependence on modality-specific channels through text-guided perturbation

Exploring Fusion Strategies for Multimodal Vision-Language Systems

Nov 26, 2025
RW
Regan Willis
🏛️ University of South Carolina

This study investigates the impact of fusion timing on the accuracy–latency trade-off in multimodal vision–language systems. We propose and systematically evaluate three fusion strategies—early, middle, and late—within a unified architecture combining BERT for language and lightweight visual backbones (MobileNetV2 or ViT) on the CMU MOSI dataset; inference latency is empirically measured on an NVIDIA Jetson Orin AGX edge platform. Results show that late fusion achieves the highest accuracy (12.3% lower MAE), while early fusion incurs the lowest latency (41.7% reduction on average), with fusion stage exhibiting a strong negative correlation between accuracy and latency. To our knowledge, this is the first work to quantitatively and systematically validate the critical influence of fusion location under a consistent experimental framework. Our findings provide reproducible architectural guidelines and empirical evidence for designing efficient multimodal models tailored to resource-constrained edge devices.

Evaluates early, intermediate, and late fusion using BERT and vision networks.Explores trade-offs between accuracy and latency in data fusion.Investigates fusion strategies for multimodal vision-language systems.

DATA: Domain-And-Time Alignment for High-Quality Feature Fusion in Collaborative Perception

Jul 24, 2025
CT
Chengchang Tian
🏛️ Southeast University | Washington State University

In collaborative perception, hardware heterogeneity induces feature-domain shift, while communication latency causes temporal misalignment—jointly degrading feature quality and accumulating cross-node errors. To address these challenges at the feature-level fusion stage, we propose a systematic alignment framework: (1) a consistency-preserving domain alignment module mitigates inter-device feature distribution discrepancies; (2) a progressive temporal alignment module corrects dynamic timing offsets via multi-scale motion modeling and two-stage compensation; and (3) an observability-constrained discriminator and instance-aware hierarchical aggregation strategy enhance semantic consistency. Evaluated on three benchmark datasets, our method achieves state-of-the-art performance and demonstrates significantly improved robustness under high communication latency and pose estimation errors.

Address domain gaps from hardware diversity and deployment conditionsEnhance semantic feature quality for collaborative perception fusionMitigate temporal misalignment caused by transmission delays

LEGO: Learnable Expansion of Graph Operators for Multi-Modal Feature Fusion

Oct 02, 2024
DD
Dexuan Ding
🏛️ Australian National University | Data61/CSIRO | Curtin University

To address weak structural modeling, shallow cross-modal interactions, difficult alignment, and poor interpretability in fusing heterogeneous multimodal features—spanning domains, granularities (e.g., token, patch, frame, clip), and modalities—this paper proposes a relation-centered, learnable graph-power fusion paradigm. It maps high-dimensional features into an interpretable graph space and constructs cross-granularity relational graphs. A learnable graph-power operator is introduced to aggregate element-wise relational scores via multivariate polynomials over homogeneous graphs, enabling structural-aware deep interaction. The method balances expressive power and interpretability, achieving multimodal fusion (text, image, video) without explicit alignment. Evaluated on video anomaly detection, it significantly outperforms concatenation, attention-based, and conventional nonlinear fusion baselines, demonstrating strong generalization and effectiveness.

Captures deep feature interactions in graph spaceImproves video anomaly detection across domainsLearnable graph fusion for multi-modal features

FusionMamba: Dynamic Feature Enhancement for Multimodal Image Fusion with Mamba

Apr 15, 2024
XX
Xinyu Xie
🏛️ Great Bay University | The Hong Kong Polytechnic University | Guangdong Institute of Intelligence Science and Technology | Macao Polytechnic University | South China University

Existing CNNs suffer from limited global modeling capacity, while ViTs incur prohibitive computational overhead—both leading to incomplete information preservation and loss of fine details in multimodal image fusion. To address these limitations, this paper proposes a dynamic feature enhancement fusion framework built upon the Mamba architecture. Our key contributions are: (1) the first visual state-space model integrating dynamic convolution with channel-wise attention; (2) a Dynamic Feature Fusion Module (DFFM) that jointly enhances texture, perceives disparity, and models cross-modal correlations; and (3) a Cross-Modal Fusion Mamba module (CMFM) to improve inter-modal interaction efficiency. Evaluated on infrared–visible image fusion and other multimodal tasks, our method achieves state-of-the-art performance, significantly improving detail richness and structural fidelity of fused images, while also boosting downstream recognition accuracy.

Global Information ProcessingImage FusionMultimodal Imaging

Latest Papers

What's happening recently
View more

Existing visual State Space Models (SSMs) rely on fixed image scanning orders, limiting their ability to model complex geometric structures and lacking effective mechanisms for multimodal interaction, which hinders their application in tasks such as multi-view 3D perception. This work proposes Deformba, the first SSM framework to incorporate deformable spatial sampling for context-adaptive spatial structure modeling, alongside a cross-attention-inspired mechanism enabling cross-modal state fusion. Deformba achieves significantly enhanced visual modeling capabilities while maintaining linear computational complexity. It provides a unified architecture for both 2D and 3D vision tasks, delivering state-of-the-art performance across multiple benchmarks—including image classification, object detection, segmentation, and bird’s-eye-view (BEV) perception—demonstrating its effectiveness and broad applicability.

3D PerceptionFixed ScanningMulti-modal Fusion

This work addresses the degradation in fusion quality arising from spatiotemporal heterogeneity in vehicular collaborative perception, caused by clock asynchrony, communication delays, and motion discrepancies. To mitigate these issues, the authors propose a dynamic compensation method that jointly models network time synchronization and Age of Information (AoI). By establishing a unified time reference and leveraging AoI to estimate communication latency, the approach enables precise spatiotemporal alignment of multi-vehicle perception features. Furthermore, it performs uncertainty-aware dynamic weighted fusion based on alignment quality and AoI. This is the first method to synergistically integrate network synchronization with AoI modeling for compensating time-varying clock drift and communication delays. Experimental results in simulated environments with clock drift and link delays demonstrate significant improvements over existing baselines, effectively enhancing the consistency and accuracy of collaborative perception.

Age of Informationclock driftcollaborative perception

This work addresses the lack of principled understanding in existing literature regarding the choice between cross-attention and feature concatenation strategies for multimodal fusion, which has largely relied on empirical heuristics. Through controlled experiments and theoretical analysis, we demonstrate for the first time that feature alignment quality is the key determinant of fusion strategy performance: under pre-aligned features, concatenation consistently outperforms cross-attention by 4.1–5.1 percentage points across all data scales, with its advantage becoming more pronounced as alignment degrades. Building on this insight, we develop a theoretical decision framework grounded in sample complexity and validate our findings using features extracted from ResNet-18 and CLIP ViT-B/32 on controlled datasets.

concatenationcross-attentionfeature alignment

Existing unified vision models primarily emphasize functional integration but lack the capacity for collaborative reasoning across image, video, and 3D modalities, hindering effective fusion of complementary priors. This work proposes PolyV, the first unified architecture enabling bidirectional interaction and mutual optimization across visual modalities. PolyV employs a sparsely gated mixture-of-experts structure with dynamic modality routing, coupled with a collaborative perception training paradigm that incorporates object- and relation-level alignment, knowledge distillation, and a coarse-to-fine collaborative fine-tuning strategy. Evaluated on ten benchmarks spanning image, video, and 3D understanding, PolyV achieves an average performance gain exceeding 10% over current state-of-the-art methods, demonstrating the efficacy of deep cross-modal collaboration.

complementary priorscross-vision synergymultimodal reasoning

This work addresses the scalability bottleneck in feed-forward novel view synthesis, where Transformer-based architectures suffer from prohibitive computational and memory overheads as the number of input views increases. To this end, we propose the AIMS framework, whose core innovation lies in decoupling the number of observations from the number of processed views via an anchor integration mechanism. Specifically, AIMS employs farthest point sampling to select a fixed set of anchors, combined with spatial clustering and a lightweight learnable integrator to aggregate information from neighboring observations. This design enables the efficient utilization of large-scale multi-view data under a constant global view budget. Experimental results demonstrate that AIMS achieves PSNR values of 29.41 dB on RealEstate10K and 17.73 dB on ScanNet, with an average rendering time of only 7.24 milliseconds per view, yielding a superior quality-efficiency trade-off.

Computational EfficiencyFeed-forwardMulti-View

Hot Scholars

MP

Marc Pollefeys

Professor of Computer Science, ETH Zurich, and Director Spatial AI Lab, Microsoft
Computer VisionComputer GraphicsRoboticsMachine Learning
GH

Gim Hee Lee

Associate Professor of Computer Science, National University of Singapore
Computer VisionRoboticsMachine Learning
KY

Kailun Yang

Professor. School of Artificial Intelligence and Robotics, Hunan University (HNU); KIT; UAH; ZJU
Computer VisionComputational OpticsIntelligent VehiclesAutonomous Driving
JX

Jin Xie

Nanjing University, China
3D Computer vision
TD

Tianchen Deng

Shanghai Jiao Tong University
RoboticsComputer Vision