Less-to-More Generalization: Unlocking More Controllability by In-Context Generation

📅 2025-04-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address weak data scalability and poor subject generalization in multi-subject image generation, this paper proposes the UNO framework. Methodologically, it first introduces an in-context multi-subject pairing data synthesis approach based on diffusion Transformers to enable zero-shot subject extension. Second, it designs the UNO model, integrating a progressive cross-modal alignment mechanism with universal rotary position encoding to support controllable single- and multi-subject joint generation. Departing from conventional fine-tuning paradigms, UNO enhances data efficiency via iterative multi-image conditional generation and in-context learning. Experiments demonstrate that UNO significantly outperforms existing methods in subject consistency, text fidelity, and layout controllability. Notably, it exhibits strong generalization capability in few-shot and even zero-shot multi-subject scenarios, establishing new state-of-the-art performance in controllable multi-subject image generation.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Generation

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGEconomics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAIUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Although subject-driven generation has been extensively explored in image generation due to its wide applications, it still has challenges in data scalability and subject expansibility. For the first challenge, moving from curating single-subject datasets to multiple-subject ones and scaling them is particularly difficult. For the second, most recent methods center on single-subject generation, making it hard to apply when dealing with multi-subject scenarios. In this study, we propose a highly-consistent data synthesis pipeline to tackle this challenge. This pipeline harnesses the intrinsic in-context generation capabilities of diffusion transformers and generates high-consistency multi-subject paired data. Additionally, we introduce UNO, which consists of progressive cross-modal alignment and universal rotary position embedding. It is a multi-image conditioned subject-to-image model iteratively trained from a text-to-image model. Extensive experiments show that our method can achieve high consistency while ensuring controllability in both single-subject and multi-subject driven generation.
Problem

Research questions and friction points this paper is trying to address.

Addressing data scalability in multi-subject image generation
Overcoming subject expansibility limits in current methods
Ensuring high consistency in multi-subject driven generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Highly-consistent data synthesis pipeline
UNO with cross-modal alignment
Diffusion transformers for multi-subject data
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shaojin Wu
Intelligent Creation Team, ByteDance
Mengqi Huang
Mengqi Huang
University of Science and Technology of China
Image GenerationVideo GenerationUnified Multimodal GenerationGenerative AI
W
Wenxu Wu
Intelligent Creation Team, ByteDance
Y
Yufeng Cheng
Intelligent Creation Team, ByteDance
Fei Ding
Fei Ding
Unknown affiliation
Qian He
Qian He
ByteDance