Score
Design and implement adaptive fusion modules that aggregate feature streams from multiple spatial or temporal views (images, frames, 3D feature maps or other per-view embeddings) into a single fused representation. These modules use learned per-view weighting and interaction mechanisms—attention, multi-head/transformer layers, multiplicative gating, latent fusion or data-driven gates—conditioned on viewpoint, distortion or state to suppress noisy inputs, amplify informative streams, and enable cross-view geometric or temporal feature interaction.
Multimodal image fusion faces two key challenges: gradient conflicts arising from cross-modal parameter sharing, which degrade performance; and modality-specific encoders that improve fusion quality yet harm task generalization. To address these, we propose a unified fusion framework integrating three novel components: semantic-aware channel pruning (to retain discriminative features), geometric affine modulation (to model inter-modal spatial discrepancies), and text-guided channel perturbation (to inject semantic priors and enhance robustness). Our method synergistically leverages pretrained semantic knowledge, channel-level perturbations, and affine transformations—achieving selective feature learning and strong cross-task generalization without introducing modality-specific parameters. Extensive experiments demonstrate consistent and significant improvements over state-of-the-art methods on major fusion benchmarks and downstream detection and segmentation tasks.
This study investigates the impact of fusion timing on the accuracy–latency trade-off in multimodal vision–language systems. We propose and systematically evaluate three fusion strategies—early, middle, and late—within a unified architecture combining BERT for language and lightweight visual backbones (MobileNetV2 or ViT) on the CMU MOSI dataset; inference latency is empirically measured on an NVIDIA Jetson Orin AGX edge platform. Results show that late fusion achieves the highest accuracy (12.3% lower MAE), while early fusion incurs the lowest latency (41.7% reduction on average), with fusion stage exhibiting a strong negative correlation between accuracy and latency. To our knowledge, this is the first work to quantitatively and systematically validate the critical influence of fusion location under a consistent experimental framework. Our findings provide reproducible architectural guidelines and empirical evidence for designing efficient multimodal models tailored to resource-constrained edge devices.
In collaborative perception, hardware heterogeneity induces feature-domain shift, while communication latency causes temporal misalignment—jointly degrading feature quality and accumulating cross-node errors. To address these challenges at the feature-level fusion stage, we propose a systematic alignment framework: (1) a consistency-preserving domain alignment module mitigates inter-device feature distribution discrepancies; (2) a progressive temporal alignment module corrects dynamic timing offsets via multi-scale motion modeling and two-stage compensation; and (3) an observability-constrained discriminator and instance-aware hierarchical aggregation strategy enhance semantic consistency. Evaluated on three benchmark datasets, our method achieves state-of-the-art performance and demonstrates significantly improved robustness under high communication latency and pose estimation errors.
To address weak structural modeling, shallow cross-modal interactions, difficult alignment, and poor interpretability in fusing heterogeneous multimodal features—spanning domains, granularities (e.g., token, patch, frame, clip), and modalities—this paper proposes a relation-centered, learnable graph-power fusion paradigm. It maps high-dimensional features into an interpretable graph space and constructs cross-granularity relational graphs. A learnable graph-power operator is introduced to aggregate element-wise relational scores via multivariate polynomials over homogeneous graphs, enabling structural-aware deep interaction. The method balances expressive power and interpretability, achieving multimodal fusion (text, image, video) without explicit alignment. Evaluated on video anomaly detection, it significantly outperforms concatenation, attention-based, and conventional nonlinear fusion baselines, demonstrating strong generalization and effectiveness.
Existing CNNs suffer from limited global modeling capacity, while ViTs incur prohibitive computational overhead—both leading to incomplete information preservation and loss of fine details in multimodal image fusion. To address these limitations, this paper proposes a dynamic feature enhancement fusion framework built upon the Mamba architecture. Our key contributions are: (1) the first visual state-space model integrating dynamic convolution with channel-wise attention; (2) a Dynamic Feature Fusion Module (DFFM) that jointly enhances texture, perceives disparity, and models cross-modal correlations; and (3) a Cross-Modal Fusion Mamba module (CMFM) to improve inter-modal interaction efficiency. Evaluated on infrared–visible image fusion and other multimodal tasks, our method achieves state-of-the-art performance, significantly improving detail richness and structural fidelity of fused images, while also boosting downstream recognition accuracy.
Existing visual State Space Models (SSMs) rely on fixed image scanning orders, limiting their ability to model complex geometric structures and lacking effective mechanisms for multimodal interaction, which hinders their application in tasks such as multi-view 3D perception. This work proposes Deformba, the first SSM framework to incorporate deformable spatial sampling for context-adaptive spatial structure modeling, alongside a cross-attention-inspired mechanism enabling cross-modal state fusion. Deformba achieves significantly enhanced visual modeling capabilities while maintaining linear computational complexity. It provides a unified architecture for both 2D and 3D vision tasks, delivering state-of-the-art performance across multiple benchmarks—including image classification, object detection, segmentation, and bird’s-eye-view (BEV) perception—demonstrating its effectiveness and broad applicability.
This work addresses the degradation in fusion quality arising from spatiotemporal heterogeneity in vehicular collaborative perception, caused by clock asynchrony, communication delays, and motion discrepancies. To mitigate these issues, the authors propose a dynamic compensation method that jointly models network time synchronization and Age of Information (AoI). By establishing a unified time reference and leveraging AoI to estimate communication latency, the approach enables precise spatiotemporal alignment of multi-vehicle perception features. Furthermore, it performs uncertainty-aware dynamic weighted fusion based on alignment quality and AoI. This is the first method to synergistically integrate network synchronization with AoI modeling for compensating time-varying clock drift and communication delays. Experimental results in simulated environments with clock drift and link delays demonstrate significant improvements over existing baselines, effectively enhancing the consistency and accuracy of collaborative perception.
This work addresses the lack of principled understanding in existing literature regarding the choice between cross-attention and feature concatenation strategies for multimodal fusion, which has largely relied on empirical heuristics. Through controlled experiments and theoretical analysis, we demonstrate for the first time that feature alignment quality is the key determinant of fusion strategy performance: under pre-aligned features, concatenation consistently outperforms cross-attention by 4.1–5.1 percentage points across all data scales, with its advantage becoming more pronounced as alignment degrades. Building on this insight, we develop a theoretical decision framework grounded in sample complexity and validate our findings using features extracted from ResNet-18 and CLIP ViT-B/32 on controlled datasets.
Existing unified vision models primarily emphasize functional integration but lack the capacity for collaborative reasoning across image, video, and 3D modalities, hindering effective fusion of complementary priors. This work proposes PolyV, the first unified architecture enabling bidirectional interaction and mutual optimization across visual modalities. PolyV employs a sparsely gated mixture-of-experts structure with dynamic modality routing, coupled with a collaborative perception training paradigm that incorporates object- and relation-level alignment, knowledge distillation, and a coarse-to-fine collaborative fine-tuning strategy. Evaluated on ten benchmarks spanning image, video, and 3D understanding, PolyV achieves an average performance gain exceeding 10% over current state-of-the-art methods, demonstrating the efficacy of deep cross-modal collaboration.
This work addresses the scalability bottleneck in feed-forward novel view synthesis, where Transformer-based architectures suffer from prohibitive computational and memory overheads as the number of input views increases. To this end, we propose the AIMS framework, whose core innovation lies in decoupling the number of observations from the number of processed views via an anchor integration mechanism. Specifically, AIMS employs farthest point sampling to select a fixed set of anchors, combined with spatial clustering and a lightweight learnable integrator to aggregate information from neighboring observations. This design enables the efficient utilization of large-scale multi-view data under a constant global view budget. Experimental results demonstrate that AIMS achieves PSNR values of 29.41 dB on RealEstate10K and 17.73 dB on ScanNet, with an average rendering time of only 7.24 milliseconds per view, yielding a superior quality-efficiency trade-off.