ORCA: Hunting Compositional Failures in Text-to-Image Diffusion

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of text-to-image diffusion models in attribute binding, spatial relationships, and counting under compositional prompts, identifying cross-modal representation space misalignment rather than information deficiency as the root cause. Accordingly, this work proposes ORCA, an Orthogonal Residual Composition Alignment method that integrates a residual parameterized predictor combining T5 and CLIP embeddings with low-rank objective alignment via a frozen visual encoder. This design seamlessly incorporates auxiliary losses into training without incurring inference overhead. Experiments demonstrate that ORCA significantly improves FID and GenEval metrics across multiple backbones, surpassing existing baselines at half the training cost while substantially enhancing generation accuracy for complex multi-object scenes.
📝 Abstract
Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count. Recent architectures already augment CLIP with a T5 encoder precisely because CLIP's contrastive embedding loses compositional structure, yet these failures persist. We argue the binding problem is therefore not one of missing information but of misaligned information: a text encoder preserves compositional structure, but in a representation space shaped by language modelling rather than vision, and the denoising objective does not directly reward aligning the two. We show this correspondence can be supplied as an explicit training signal, that the relevant cross-modal information is concentrated in a low-rank subspace of self-supervised visual features, and that supplying it can be folded into diffusion training as a single auxiliary loss. Our method, ORCA (Orthogonal Residual Compositional Alignment), aligns the latent of a diffusion transformer with a low-rank target derived from a frozen visual encoder, through a predictor whose orthogonal basis is parameterised by a learned residual between T5 and CLIP embeddings, which provides a prompt-dependent signal for selecting the visual readout subspace. We prove that the cross-modal information recoverable at a given rank is bounded by the spectral mass of the visual encoder's covariance in the top components. Across three diffusion-transformer backbones (DiT-B/2, DiT-L/2, U-ViT-L), ORCA improves FID and GenEval over both vanilla and REPA baselines at zero inference-time cost; on DiT-L/2 it reaches FID 16.65 and GenEval 0.291 at 200K steps, exceeding the strongest 400K baseline at half the training cost, with the largest gains concentrated on attribute binding, spatial relations, and multi-object prompts.
Problem

Research questions and friction points this paper is trying to address.

text-to-image diffusion
compositional generation
binding problem
cross-modal alignment
compositional failures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compositional Alignment
Low-rank Subspace
Diffusion Transformer
Orthogonal Residual
Cross-modal Representation
A
Arshia Hemmat
Wellcome Sanger Institute; Cambridge Stem Cell Institute, University of Cambridge; Cambridge Centre for AI in Medicine, University of Cambridge
A
Amirhossein Vahidi
Wellcome Sanger Institute; Cambridge Centre for AI in Medicine, University of Cambridge
Amitis Shidani
Amitis Shidani
University of Oxford
Applied StatisticsMachine LearningRepresentation LearningLearning TheoryOptimization
M
Mohammad Vali Sanian
Wellcome Sanger Institute; Department of Computer Science, University of Helsinki; Institute for Molecular Medicine Finland (FIMM), University of Helsinki
H
Hesam Asadollahzadeh
Wellcome Sanger Institute; School of Computing and Information Systems, University of Melbourne
A
Aryan Yazdan Parast
School of Computing and Information Systems, University of Melbourne
M
Mohammad Lotfollahi
Wellcome Sanger Institute; Cambridge Stem Cell Institute, University of Cambridge; Cambridge Centre for AI in Medicine, University of Cambridge; Department of Medicine, University of Cambridge