Where to Look Is Not How to Fix: Pre-Denoising Diagnostics and Modality-Dependent Control in Diffusion Composition

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the mismatch between diagnostic localization and intervention efficacy in compositional generation with diffusion models. We propose the "Anchor-Stress Protocol" and a text-only Compositional Stress Index (CSI) to jointly evaluate text encoder diagnostics and denoiser interventions. Through modular analysis on Stable Diffusion architectures—employing embedding adapters, cross-attention interventions, and selective amplification/subtraction techniques—we reveal a "diagnostic-control decoupling" phenomenon: feature coordinates exposing risks are not necessarily optimal intervention points. Furthermore, our findings demonstrate that deeper encoder layers facilitate defect diagnosis, whereas decoder blocks substantially improve color accuracy, thereby elucidating architecture-dependent response mechanisms.
📝 Abstract
Understanding compositional failures in text-to-image diffusion requires identifying both where stress is detectable and how intervention changes the output. We study these questions through a controlled anchor--stress protocol that jointly evaluates text-encoder diagnostics and denoiser interventions. We introduce a text-only Compositional Stress Index (CSI), which separates common from rare compositions across SD1.5, SDXL, and the SD3 text path and provides an upstream diagnostic coordinate. A matched six-prompt localization study links intervention location to distinct outcomes: residual-minimizing embedding adapters improve representation fit, while downstream cross-attention intervention increases color hit rate (CHR) by 0.0272. Across SD1.5 and SDXL denoiser blocks, the largest positive signed diagnostic-accessibility mean occurs at the deep encoder, whereas selective boost has its largest positive mean CHR response at decoder blocks. Selective subtraction and broad ablation reveal further modality- and architecture-dependent responses, including a substantial CHR decrease when SDXL decoder cross-attention is broadly ablated. We find a diagnosis-control dissociation under our controlled attribute-object composition setting: compositional defects are diagnosable before denoising, but the representation coordinate that exposes risk is not necessarily the coordinate or modality that improves generation.
Problem

Research questions and friction points this paper is trying to address.

diffusion models
compositional generation
text-to-image
diagnosis-control dissociation
compositional stress
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compositional Stress Index
Diffusion Composition
Diagnosis-Control Dissociation
Cross-Attention Intervention
Text-to-Image Diagnostics
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
F
Fangzheng Wu
Department of Computer Science, Tulane University, New Orleans, LA 70118
Brian Summa
Brian Summa
Associate Professor, Tulane University