Unleashing the Power of Text: Text-Guided Flow Matching for Image Fusion under Complex Degradations

📅 2026-08-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of unreliable modality information in infrared and visible image fusion under real-world complex degradation scenarios by proposing TGFusion, a novel framework that encodes task specifications, degradation characteristics, and generative intent into structured textual prompts. Within a latent space, it unifies degradation suppression and cross-modal fusion through flow matching. The framework introduces a Prompt-conditioned Multi-Stream Joint Flow Transformer, treating text as an independent semantic stream processed in parallel with visual streams; bidirectional token-level interactions between these streams are enabled via joint attention to dynamically guide information selection and fused image generation. Extensive experiments demonstrate that TGFusion achieves superior performance across multiple benchmarks and complex degradation settings, excelling in perceptual quality, naturalness, structural detail preservation, and infrared saliency retention, while exhibiting strong robustness against diverse single and composite degradations.
📝 Abstract
Infrared-visible image fusion under realistic degradation scenarios is a challenging task, as degradations not only cause a loss of reliable modality-specific information in observed images but also hinder the fusion process. Recent studies indicate that text can provide prior information about degradation characteristics, complementing the limited evidence available from corrupted input images and facilitating fusion. However, existing methods typically inject fixed global text representations into visual features, making it difficult for textual guidance to adapt to spatially varying degradations, local structures, and thermal saliency. To this end, we propose TGFusion, a text-guided latent-space flow matching framework that unifies degradation suppression and cross-modal fusion. TGFusion encodes task, degradation, and generation cues into structured prompts. To fully exploit these priors, we design a Prompt-conditioned Multi-stream Joint Flow Transformer that represents text as an independent semantic stream alongside fusion, visible, and infrared streams. Joint attention enables token-level bidirectional interaction and layer-wise updating among semantic and visual representations, allowing degradation semantics to dynamically guide reliable information selection and fusion latent generation. Extensive experiments on public benchmarks and complex degradation scenarios demonstrate that TGFusion achieves superior or competitive performance in perceptual quality, image naturalness, structural-detail preservation, and infrared-saliency retention, while remaining robust across diverse single and compound degradations.
Problem

Research questions and friction points this paper is trying to address.

infrared-visible image fusion
complex degradations
text-guided fusion
degradation suppression
cross-modal fusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

text-guided fusion
flow matching
multi-stream transformer
degradation-aware prompting
infrared-visible image fusion
🔎 Similar Papers
No similar papers found.