text-driven image fusion

Designs and builds image-fusion models and pipelines that combine multiple visual inputs into a single output under control of natural-language prompts, implementing language-conditioning mechanisms (e.g., dual-attention, text encoders, attention-based fusion modules) to steer which visual attributes are preserved or merged. Analyzes and evaluates the semantic alignment between fused outputs and text priors and the adaptation of fusion behavior to varying input content and prompts.

text-drivenimagefusion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.52
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

FUSION: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding

Apr 14, 2025
ZL
Zheng Liu
🏛️ Peking University | Shanghai AI Laboratory | Nanjing University

This study addresses insufficient deep integration of visual and linguistic representations. We propose FUSION-3B, a multimodal large language model featuring full-modality dynamic alignment. Methodologically, we introduce three novel components: (1) text-guided unified visual encoding, (2) context-aware recursive alignment decoding, and (3) dual-supervised semantic mapping loss—collectively transcending conventional late-fusion paradigms to enable pixel-level, question-level, and end-to-end cross-modal unified modeling. Despite its compact 3B parameter count, FUSION-3B achieves superior performance on most benchmarks using only 630 visual tokens—outperforming Cambrian-1 (8B) and Florence-VL (8B); even with only 300 visual tokens, it surpasses Cambrian-1 (8B), and it exceeds LLaVA-NeXT on over half of the evaluated benchmarks. To further support fine-grained vision–language alignment, we construct a language-driven synthetic QA dataset.

Achieves deep vision-language integration in MLLMsEnables fine-grained semantic alignment during decodingMitigates modality discrepancies with supervised loss

This study addresses the unclear mechanisms underlying cross-layer fusion of visual and textual information in multimodal large language models. We propose an architecture-aware diagnostic framework that systematically compares concatenation-based and native multimodal architectures. Through alignment decoupling, attention entropy analysis, intrinsic dimensionality estimation, causal intervention, and vision-specific Centered Kernel Alignment (CKA), we reveal how feature spaces are reorganized under different architectural paradigms. Our findings indicate that concatenation-based models exhibit a text-dominant fusion trajectory, whereas native models achieve early-stage vision-language co-adaptation. This work elucidates the distinct multimodal fusion mechanisms inherent to these two architectural paradigms, providing a principled theoretical foundation for future model design.

Architecture ParadigmsMechanistic InterpretabilityMultimodal Fusion

This work addresses the lack of principled understanding in existing literature regarding the choice between cross-attention and feature concatenation strategies for multimodal fusion, which has largely relied on empirical heuristics. Through controlled experiments and theoretical analysis, we demonstrate for the first time that feature alignment quality is the key determinant of fusion strategy performance: under pre-aligned features, concatenation consistently outperforms cross-attention by 4.1–5.1 percentage points across all data scales, with its advantage becoming more pronounced as alignment degrades. Building on this insight, we develop a theoretical decision framework grounded in sample complexity and validate our findings using features extracted from ResNet-18 and CLIP ViT-B/32 on controlled datasets.

concatenationcross-attentionfeature alignment

Multimodal Alignment and Fusion: A Survey

Nov 26, 2024
SL
Songtao Li
🏛️ Northeastern University | Peking University

This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.

Addressing cross-modal misalignment and computational bottlenecks challengesExploring applications from social media to medical imagingSurveying multimodal alignment and fusion techniques in machine learning

This work challenges the prevailing view that image and text representations in vision-language models align only in deep layers. Inspired by DeepDream, the authors propose a synthesis method within an adapter-based architecture that extracts textual concept vectors layer by layer and optimizes corresponding images without auxiliary models or additional datasets. For the first time, this approach provides direct, concept-level evidence of cross-modal alignment starting from the very first layer: across seven network layers and hundreds of concepts, over 50% of images synthesized from the first layer already exhibit clear, salient visual features of target concepts—such as animals, activities, and seasons. This method offers an efficient and novel pathway toward enhancing the interpretability of vision-language models.

concept alignmentimage-text alignmentmultimodal representation

Latest Papers

What's happening recently
View more

该研究提出了一种先对齐后融合的框架,通过三重配对余弦对齐和提示引导查询解码器来解决3D视觉-语言任务中的特征不一致问题。

3D vision-languagefeature discrepancyheterogeneous representations

Exploring Fusion Strategies for Multimodal Vision-Language Systems

Nov 26, 2025
RW
Regan Willis
🏛️ University of South Carolina

This study investigates the impact of fusion timing on the accuracy–latency trade-off in multimodal vision–language systems. We propose and systematically evaluate three fusion strategies—early, middle, and late—within a unified architecture combining BERT for language and lightweight visual backbones (MobileNetV2 or ViT) on the CMU MOSI dataset; inference latency is empirically measured on an NVIDIA Jetson Orin AGX edge platform. Results show that late fusion achieves the highest accuracy (12.3% lower MAE), while early fusion incurs the lowest latency (41.7% reduction on average), with fusion stage exhibiting a strong negative correlation between accuracy and latency. To our knowledge, this is the first work to quantitatively and systematically validate the critical influence of fusion location under a consistent experimental framework. Our findings provide reproducible architectural guidelines and empirical evidence for designing efficient multimodal models tailored to resource-constrained edge devices.

Evaluates early, intermediate, and late fusion using BERT and vision networks.Explores trade-offs between accuracy and latency in data fusion.Investigates fusion strategies for multimodal vision-language systems.

This work addresses the unclear mechanisms of vision–language integration in current multimodal large language models (MLLMs). Through layer-wise masking analysis and attention evolution tracking, the study systematically reveals for the first time that cross-modal fusion predominantly occurs in specific layers and identifies a late-stage “retrospective” reactivation of visual signals. Building on these insights, the authors propose a training-free contrastive attention framework that guides the model to enhance meaningful cross-modal attention transfer. Extensive experiments across diverse mainstream MLLM architectures and multimodal benchmarks demonstrate the effectiveness of the proposed mechanism, yielding significant improvements in multimodal reasoning performance.

Attention MechanismLayer-wise AnalysisMultimodal Large Language Models

Multi-Grained Text-Guided Image Fusion for Multi-Exposure and Multi-Focus Scenarios

Dec 23, 2025
MT
Mingwei Tang
🏛️ Xidian University | Nanyang Technological University | Xi'an University of Technology

To address the challenge of jointly modeling dynamic range and depth-of-field variations in multi-exposure and multi-focus image fusion, this paper proposes the first hierarchical text-guided fusion framework. Methodologically, we design a multi-granularity text encoder—capturing fine-grained details, mid-granularity structures, and coarse-granularity semantics—and build a hierarchical cross-modulation network. We further introduce a granularity-aware supervised loss and a saliency-driven semantic enhancement module to achieve precise cross-modal feature alignment and adaptive modulation. Our key innovation lies in explicitly embedding text granularity priors into the fusion process during training, eliminating the need for textual input at test time. Extensive experiments on mainstream multi-exposure and multi-focus benchmarks demonstrate consistent superiority over state-of-the-art methods, with significant improvements in PSNR and SSIM. The framework exhibits strong generalization across diverse fusion scenarios.

Addresses disparities in dynamic range and focus depth between imagesEnhances fusion with multi-grained text guidance and cross-modal alignmentSynthesizes high-quality images from varied exposure and focus inputs

This work addresses the limitation of existing vision prompt tuning methods, which rely on a single image-prompt fusion strategy. The authors formulate the selection of fusion mechanisms as a bilevel optimization problem and employ differentiable architecture search to jointly optimize prompts and their layer-specific fusion schemes across Transformer layers. They introduce novel fusion operations combining affine transformations and cross-attention, enabling the first automatic discovery of heterogeneous, layer-adaptive fusion strategies. Extensive experiments across 34 datasets demonstrate that the proposed method significantly outperforms current prompt tuning baselines when using a frozen ViT backbone, achieving a superior trade-off among accuracy, inference latency, and parameter efficiency.

fusion schemelayer semanticsprompt fusion

Hot Scholars

JZ

Jinchao Zhang

WeChat AI - Pattern Recognition Center
Deep LearningNatural Language ProcessingMachine TranslationDialogue System
JL

Jiasen Lu

Research Scientist, Apple
Computer VisionNatural Language Processing
WY

Wenming Yang

Tsinghua University
Computer VisionImage Processing
CL

Chun-Liang Li

Apple / University of Washington
Machine LearningStatistics