DLSF: Dual-Layer Synergistic Fusion for High-Fidelity Image Syn-thesis

📅 2025-07-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Stable Diffusion models often suffer from incomplete semantic alignment and fine-grained texture loss in complex scenes due to insufficient feature aggregation capacity. To address this, we propose a two-level collaborative latent fusion framework comprising an Adaptive Global Fusion (AGF) module and a Dynamic Spatial Fusion (DSF) module, which jointly model cross-level feature interactions between a base layer and a refinement layer. We further introduce hierarchical coordination and spatially aware refinement mechanisms to enhance structural coherence and detail fidelity. Crucially, our approach achieves these improvements without increasing inference overhead. Extensive evaluations on multiple benchmarks demonstrate that our method outperforms state-of-the-art diffusion models—particularly in high-texture, multi-object complex scenes—delivering superior structural preservation and fine-detail recovery while maintaining global semantic consistency and local texture fidelity.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Sentence-level Semantics, Textual Inference, etc.

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAG
📝 Abstract
With the rapid advancement of diffusion-based generative models, Stable Diffusion (SD) has emerged as a state-of-the-art framework for high-fidelity im-age synthesis. However, existing SD models suffer from suboptimal feature aggregation, leading to in-complete semantic alignment and loss of fine-grained details, especially in highly textured and complex scenes. To address these limitations, we propose a novel dual-latent integration framework that en-hances feature interactions between the base latent and refined latent representations. Our approach em-ploys a feature concatenation strategy followed by an adaptive fusion module, which can be instantiated as either (i) an Adaptive Global Fusion (AGF) for hier-archical feature harmonization, or (ii) a Dynamic Spatial Fusion (DSF) for spatially-aware refinement. This design enables more effective cross-latent com-munication, preserving both global coherence and local texture fidelity. Our GitHub page: https://anonymous.4open.science/r/MVA2025-22 .
Problem

Research questions and friction points this paper is trying to address.

Suboptimal feature aggregation in Stable Diffusion models
Incomplete semantic alignment in generated images
Loss of fine-grained details in complex scenes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-latent integration enhances feature interactions
Adaptive fusion module harmonizes hierarchical features
Dynamic Spatial Fusion refines spatial awareness
🔎 Similar Papers
No similar papers found.
Z
Zhen-Qi Chen
National Yang Ming Chiao Tung University, 1001 University Road, Hsinchu 300, Taiwan
Yuan-Fu Yang
Yuan-Fu Yang
Assistant Professor, National Chiao Tung University
GenAI3D representationComputer Vision