Score
Designs and builds image-fusion models and pipelines that combine multiple visual inputs into a single output under control of natural-language prompts, implementing language-conditioning mechanisms (e.g., dual-attention, text encoders, attention-based fusion modules) to steer which visual attributes are preserved or merged. Analyzes and evaluates the semantic alignment between fused outputs and text priors and the adaptation of fusion behavior to varying input content and prompts.
This study addresses insufficient deep integration of visual and linguistic representations. We propose FUSION-3B, a multimodal large language model featuring full-modality dynamic alignment. Methodologically, we introduce three novel components: (1) text-guided unified visual encoding, (2) context-aware recursive alignment decoding, and (3) dual-supervised semantic mapping loss—collectively transcending conventional late-fusion paradigms to enable pixel-level, question-level, and end-to-end cross-modal unified modeling. Despite its compact 3B parameter count, FUSION-3B achieves superior performance on most benchmarks using only 630 visual tokens—outperforming Cambrian-1 (8B) and Florence-VL (8B); even with only 300 visual tokens, it surpasses Cambrian-1 (8B), and it exceeds LLaVA-NeXT on over half of the evaluated benchmarks. To further support fine-grained vision–language alignment, we construct a language-driven synthetic QA dataset.
This study addresses the unclear mechanisms underlying cross-layer fusion of visual and textual information in multimodal large language models. We propose an architecture-aware diagnostic framework that systematically compares concatenation-based and native multimodal architectures. Through alignment decoupling, attention entropy analysis, intrinsic dimensionality estimation, causal intervention, and vision-specific Centered Kernel Alignment (CKA), we reveal how feature spaces are reorganized under different architectural paradigms. Our findings indicate that concatenation-based models exhibit a text-dominant fusion trajectory, whereas native models achieve early-stage vision-language co-adaptation. This work elucidates the distinct multimodal fusion mechanisms inherent to these two architectural paradigms, providing a principled theoretical foundation for future model design.
This work addresses the lack of principled understanding in existing literature regarding the choice between cross-attention and feature concatenation strategies for multimodal fusion, which has largely relied on empirical heuristics. Through controlled experiments and theoretical analysis, we demonstrate for the first time that feature alignment quality is the key determinant of fusion strategy performance: under pre-aligned features, concatenation consistently outperforms cross-attention by 4.1–5.1 percentage points across all data scales, with its advantage becoming more pronounced as alignment degrades. Building on this insight, we develop a theoretical decision framework grounded in sample complexity and validate our findings using features extracted from ResNet-18 and CLIP ViT-B/32 on controlled datasets.
This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.
This work challenges the prevailing view that image and text representations in vision-language models align only in deep layers. Inspired by DeepDream, the authors propose a synthesis method within an adapter-based architecture that extracts textual concept vectors layer by layer and optimizes corresponding images without auxiliary models or additional datasets. For the first time, this approach provides direct, concept-level evidence of cross-modal alignment starting from the very first layer: across seven network layers and hundreds of concepts, over 50% of images synthesized from the first layer already exhibit clear, salient visual features of target concepts—such as animals, activities, and seasons. This method offers an efficient and novel pathway toward enhancing the interpretability of vision-language models.
该研究提出了一种先对齐后融合的框架,通过三重配对余弦对齐和提示引导查询解码器来解决3D视觉-语言任务中的特征不一致问题。
This study investigates the impact of fusion timing on the accuracy–latency trade-off in multimodal vision–language systems. We propose and systematically evaluate three fusion strategies—early, middle, and late—within a unified architecture combining BERT for language and lightweight visual backbones (MobileNetV2 or ViT) on the CMU MOSI dataset; inference latency is empirically measured on an NVIDIA Jetson Orin AGX edge platform. Results show that late fusion achieves the highest accuracy (12.3% lower MAE), while early fusion incurs the lowest latency (41.7% reduction on average), with fusion stage exhibiting a strong negative correlation between accuracy and latency. To our knowledge, this is the first work to quantitatively and systematically validate the critical influence of fusion location under a consistent experimental framework. Our findings provide reproducible architectural guidelines and empirical evidence for designing efficient multimodal models tailored to resource-constrained edge devices.
This work addresses the unclear mechanisms of vision–language integration in current multimodal large language models (MLLMs). Through layer-wise masking analysis and attention evolution tracking, the study systematically reveals for the first time that cross-modal fusion predominantly occurs in specific layers and identifies a late-stage “retrospective” reactivation of visual signals. Building on these insights, the authors propose a training-free contrastive attention framework that guides the model to enhance meaningful cross-modal attention transfer. Extensive experiments across diverse mainstream MLLM architectures and multimodal benchmarks demonstrate the effectiveness of the proposed mechanism, yielding significant improvements in multimodal reasoning performance.
To address the challenge of jointly modeling dynamic range and depth-of-field variations in multi-exposure and multi-focus image fusion, this paper proposes the first hierarchical text-guided fusion framework. Methodologically, we design a multi-granularity text encoder—capturing fine-grained details, mid-granularity structures, and coarse-granularity semantics—and build a hierarchical cross-modulation network. We further introduce a granularity-aware supervised loss and a saliency-driven semantic enhancement module to achieve precise cross-modal feature alignment and adaptive modulation. Our key innovation lies in explicitly embedding text granularity priors into the fusion process during training, eliminating the need for textual input at test time. Extensive experiments on mainstream multi-exposure and multi-focus benchmarks demonstrate consistent superiority over state-of-the-art methods, with significant improvements in PSNR and SSIM. The framework exhibits strong generalization across diverse fusion scenarios.
This work addresses the limitation of existing vision prompt tuning methods, which rely on a single image-prompt fusion strategy. The authors formulate the selection of fusion mechanisms as a bilevel optimization problem and employ differentiable architecture search to jointly optimize prompts and their layer-specific fusion schemes across Transformer layers. They introduce novel fusion operations combining affine transformations and cross-attention, enabling the first automatic discovery of heterogeneous, layer-adaptive fusion strategies. Extensive experiments across 34 datasets demonstrate that the proposed method significantly outperforms current prompt tuning baselines when using a frozen ViT backbone, achieving a superior trade-off among accuracy, inference latency, and parameter efficiency.