RefRoute: Decoupling Conditioning Cost from References via Compact Residual Conditioning and Spatial Routing

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational overhead and poor scalability caused by dense visual tokens and global attention in multi-reference image generation. To this end, it proposes compact residual conditioning and spatial routing mechanisms. Methodologically, low-resolution latent tokens are combined with full-resolution pixel residual features to preserve fine details while compressing representations. Furthermore, conditional and attention routing restricts redundant interactions, effectively decoupling reference encoding costs from attention overhead. The authors also introduce ManyRef100, a dedicated evaluation benchmark for this task. Experimental results demonstrate that the proposed method achieves a weighted score of 36.06, significantly outperforming the FLUX baseline. Notably, in 16-reference scenarios, it yields up to an 18.3× inference speedup, effectively suppressing latency growth.
📝 Abstract
Multi-reference image generation requires preserving the appearance of multiple subjects while composing them into a coherent scene. However, existing diffusion transformers commonly encode references as dense visual token grids and jointly process them with global attention, making conditioning increasingly expensive as the number and resolution of references grow. We present RefRoute, a framework that addresses both reference representation cost and attention overhead through two complementary mechanisms. Compact residual conditioning combines low-resolution latent tokens with lightweight residual features extracted from full-resolution pixels, reducing reference token counts while retaining fine-grained appearance cues. Condition routing and attention routing align reference tokens with their assigned target regions and restrict cross-reference interactions, while allowing selective reference access beyond region boundaries for scene integration. We further introduce RefRoute-Data for training many-reference generation models and ManyRef100, a benchmark spanning human, object, and mixed compositions with 10-17 references. After many-reference fine-tuning, RefRoute achieves an overall Weighted-Ref-VIEScore of 36.06 on ManyRef100, compared with 8.88 for FLUX.2-Klein-9B. Separate inference-cost evaluations show substantially slower latency growth as the reference count increases: at 16 references, our 50-step and 4-step configurations achieve $18.3\times$ and $14.2\times$ speedups over their corresponding FLUX baselines, respectively. These results establish compact reference representations and spatially routed attention as an effective approach to scalable many-reference image generation.
Problem

Research questions and friction points this paper is trying to address.

multi-reference image generation
diffusion transformers
conditioning cost
attention overhead
scalable generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-reference image generation
Compact residual conditioning
Spatial routing
Diffusion transformers
Scalable attention