Augmented Conditioning Is Enough For Effective Training Image Generation

📅 2025-02-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address insufficient generative image diversity—which hinders downstream classification model training—this paper introduces the “Augmented Conditioning” paradigm. Without fine-tuning pre-trained text-to-image diffusion models, it jointly conditions generation on both standard-augmented real images (e.g., cropping, color jittering) and textual prompts. This approach significantly enhances visual diversity of generated samples while preserving semantic fidelity and domain consistency. As a lightweight, inference-time, zero-shot, parameter-free strategy, it achieves state-of-the-art performance across five long-tailed and extreme few-shot (1-shot/5-shot) classification benchmarks. It improves average accuracy by 3.2% on long-tailed tasks and up to 18.7% in few-shot settings, with particularly notable gains in generalization to tail and rare classes.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Generation

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Image generation abilities of text-to-image diffusion models have significantly advanced, yielding highly photo-realistic images from descriptive text and increasing the viability of leveraging synthetic images to train computer vision models. To serve as effective training data, generated images must be highly realistic while also sufficiently diverse within the support of the target data distribution. Yet, state-of-the-art conditional image generation models have been primarily optimized for creative applications, prioritizing image realism and prompt adherence over conditional diversity. In this paper, we investigate how to improve the diversity of generated images with the goal of increasing their effectiveness to train downstream image classification models, without fine-tuning the image generation model. We find that conditioning the generation process on an augmented real image and text prompt produces generations that serve as effective synthetic datasets for downstream training. Conditioning on real training images contextualizes the generation process to produce images that are in-domain with the real image distribution, while data augmentations introduce visual diversity that improves the performance of the downstream classifier. We validate augmentation-conditioning on a total of five established long-tail and few-shot image classification benchmarks and show that leveraging augmentations to condition the generation process results in consistent improvements over the state-of-the-art on the long-tailed benchmark and remarkable gains in extreme few-shot regimes of the remaining four benchmarks. These results constitute an important step towards effectively leveraging synthetic data for downstream training.
Problem

Research questions and friction points this paper is trying to address.

Enhancing diversity in text-to-image generation models
Improving synthetic image effectiveness for training
Augmenting conditioning for better classification benchmarks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Augmented real image conditioning
Text prompt integration
No fine-tuning required
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.