🤖 AI Summary
This study addresses the linear growth of memory consumption and subject identity loss in multi-reference image generation by proposing a pixel-space composition framework. Methodologically, it employs a fixed-resolution canvas to decouple layout from appearance, thereby maintaining a constant conditioning size. Furthermore, a dual-stream LoRA adapter is introduced to disentangle geometric structure from local texture, which, combined with a diffusion Transformer and depth map guidance, enables precise compositional control. Experimental results demonstrate that the proposed method outperforms existing state-of-the-art models in color, texture, shape, and identity fidelity. Notably, inference memory usage remains constant regardless of the number of references. This work thus presents an efficient solution for cinematic-grade multi-subject synthesis.
📝 Abstract
Filmmakers and visual artists routinely need to compose multiple references, actors, locations, props, cultural elements, into a single coherent shot, but existing tools either fail to scale past a handful of references or destroy fine grained subject identity in the process, since per reference tokenization scales memory linearly with reference count $N$ and generated content often departs from the given references rather than reproducing them. We propose \textbf{RefCompose}, a pixel space compositional conditioning framework that decouples \emph{where} things go from \emph{what} they look like, via a single fixed resolution reference canvas that keeps conditioning size constant regardless of reference count. Spatial layout is induced at inference time from a frozen diffusion transformer and extracted via Grounding DINO, requiring no LLM or dedicated layout model, while dual stream LoRA adapters inject a layout derived depth map and the encoded canvas through separate low rank streams, disentangling geometric scaffolding from localized appearance. On the Dense Layout protocol, RefCompose consistently outperforms layout based and state of the art multi reference baselines on color, texture, shape, spatial accuracy, and identity/content preservation at higher reference counts, all with constant inference memory, making it a practical building block for multi subject cinematic composition at production scale.