🤖 AI Summary
Stable Diffusion models often suffer from incomplete semantic alignment and fine-grained texture loss in complex scenes due to insufficient feature aggregation capacity. To address this, we propose a two-level collaborative latent fusion framework comprising an Adaptive Global Fusion (AGF) module and a Dynamic Spatial Fusion (DSF) module, which jointly model cross-level feature interactions between a base layer and a refinement layer. We further introduce hierarchical coordination and spatially aware refinement mechanisms to enhance structural coherence and detail fidelity. Crucially, our approach achieves these improvements without increasing inference overhead. Extensive evaluations on multiple benchmarks demonstrate that our method outperforms state-of-the-art diffusion models—particularly in high-texture, multi-object complex scenes—delivering superior structural preservation and fine-detail recovery while maintaining global semantic consistency and local texture fidelity.
📝 Abstract
With the rapid advancement of diffusion-based generative models, Stable Diffusion (SD) has emerged as a state-of-the-art framework for high-fidelity im-age synthesis. However, existing SD models suffer from suboptimal feature aggregation, leading to in-complete semantic alignment and loss of fine-grained details, especially in highly textured and complex scenes. To address these limitations, we propose a novel dual-latent integration framework that en-hances feature interactions between the base latent and refined latent representations. Our approach em-ploys a feature concatenation strategy followed by an adaptive fusion module, which can be instantiated as either (i) an Adaptive Global Fusion (AGF) for hier-archical feature harmonization, or (ii) a Dynamic Spatial Fusion (DSF) for spatially-aware refinement. This design enables more effective cross-latent com-munication, preserving both global coherence and local texture fidelity. Our GitHub page: https://anonymous.4open.science/r/MVA2025-22 .