RefCompose: Multi-Reference Image Generation via LoRA-Conditioned Diffusion

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the linear growth of memory consumption and subject identity loss in multi-reference image generation by proposing a pixel-space composition framework. Methodologically, it employs a fixed-resolution canvas to decouple layout from appearance, thereby maintaining a constant conditioning size. Furthermore, a dual-stream LoRA adapter is introduced to disentangle geometric structure from local texture, which, combined with a diffusion Transformer and depth map guidance, enables precise compositional control. Experimental results demonstrate that the proposed method outperforms existing state-of-the-art models in color, texture, shape, and identity fidelity. Notably, inference memory usage remains constant regardless of the number of references. This work thus presents an efficient solution for cinematic-grade multi-subject synthesis.
📝 Abstract
Filmmakers and visual artists routinely need to compose multiple references, actors, locations, props, cultural elements, into a single coherent shot, but existing tools either fail to scale past a handful of references or destroy fine grained subject identity in the process, since per reference tokenization scales memory linearly with reference count $N$ and generated content often departs from the given references rather than reproducing them. We propose \textbf{RefCompose}, a pixel space compositional conditioning framework that decouples \emph{where} things go from \emph{what} they look like, via a single fixed resolution reference canvas that keeps conditioning size constant regardless of reference count. Spatial layout is induced at inference time from a frozen diffusion transformer and extracted via Grounding DINO, requiring no LLM or dedicated layout model, while dual stream LoRA adapters inject a layout derived depth map and the encoded canvas through separate low rank streams, disentangling geometric scaffolding from localized appearance. On the Dense Layout protocol, RefCompose consistently outperforms layout based and state of the art multi reference baselines on color, texture, shape, spatial accuracy, and identity/content preservation at higher reference counts, all with constant inference memory, making it a practical building block for multi subject cinematic composition at production scale.
Problem

Research questions and friction points this paper is trying to address.

multi-reference image generation
subject identity preservation
compositional conditioning
memory scalability
cinematic composition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Reference Image Generation
LoRA-Conditioned Diffusion
Pixel-Space Compositional Conditioning
Dual-Stream LoRA Adapters
Spatial Layout Disentanglement
🔎 Similar Papers
No similar papers found.
S
Sai Sri Teja Kuppa
Eros Innovation, India
P
Parth Shinde
Eros Innovation, India; Indian Institute of Technology Madras, India
P
Priyadharsan Balaji S
Eros Innovation, India; Indian Institute of Technology Madras, India
J
Jinka Harshavardhan
Eros Innovation, India
Sriprabha Ramanarayanan
Sriprabha Ramanarayanan
Indian Institute of Technology Madras, HTIC
Medical image analysisMachine learningDeep neural networkscomputer vision