retrieval-augmented image generation

Design, build, or evaluate image synthesis systems that condition generation or editing on externally retrieved information. This includes autonomously planning and executing searches, retrieving out‑of‑distribution facts or assets, and merging retrieved results into generation prompts or edit constraints so produced images are grounded in the retrieved knowledge.

retrieval-augmentedimagegeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitation of existing image generation models, which rely on static internal knowledge and struggle to incorporate external or up-to-date information required in real-world scenarios. The authors propose the first trainable search-augmented image generation agent that leverages multi-hop retrieval to acquire both textual knowledge and reference images, enabling knowledge-grounded image synthesis. To support this approach, they introduce dedicated datasets—Gen-Searcher-SFT-10k and Gen-Searcher-RL-6k—and a new evaluation benchmark, KnowGen. They further design a dual-modality reward mechanism for reinforcement learning, combining supervised fine-tuning with the GRPO algorithm. Experimental results demonstrate substantial improvements, with performance gains of approximately 16 points on KnowGen and 15 points on the WISE benchmark, significantly outperforming strong baselines such as Qwen-Image.

external knowledgegrounded generationimage generation

Current image generation models exhibit limitations in handling ambiguous intents, logical reasoning, and out-of-distribution knowledge, primarily due to the absence of deep reasoning capabilities and real-time access to external information. This work proposes the first training-free, plug-and-play multi-stage agent framework that seamlessly integrates retrieval-augmented generation with deep reasoning. By employing a closed-loop “question-answering” mechanism, the framework autonomously identifies gaps in logic and knowledge, dynamically plans retrieval and reasoning steps, and orchestrates them in an adaptive manner. Introducing the agent paradigm into image generation for the first time, the method achieves significant performance gains on the WISE Verified and RISEBench benchmarks, delivering absolute improvements of 0.313 and 19.70 for Qwen-Image and Qwen-Image-Edit-2511, respectively, thereby establishing a new state of the art among open-source models.

ambiguous intentionsimage generationlogical reasoning

Current visual generative models exhibit significant limitations in spatial reasoning, state persistence, long-term consistency, and causal understanding, hindering their ability to produce structurally coherent and intelligently behaving content. This work proposes a paradigm shift from appearance-based synthesis toward intelligent visual generation, introducing a novel five-level generative capability taxonomy—from atomic generation to world modeling—that emphasizes the integration of structure, dynamics, domain knowledge, and causality. By leveraging key technical components including a unified understanding-generation architecture, flow matching, enhanced representations, post-training optimization, and synthetic data distillation, the study establishes a capability-centered evaluation framework. This framework exposes the prevailing overreliance on perceptual quality metrics while neglecting structural and causal deficiencies, thereby charting a roadmap for the development of next-generation intelligent visual generation systems.

causal understandingintelligent visual generationlong-horizon consistency

A Picture is Worth a Thousand Prompts? Efficacy of Iterative Human-Driven Prompt Refinement in Image Regeneration Tasks

Apr 29, 2025
KT
Khoi Trinh
🏛️ University of Oklahoma | University of Texas at San Antonio

This study investigates whether humans can accurately regenerate target images via iterative prompt optimization in AI-based image generation, and systematically evaluates the reliability of mainstream image similarity metrics (LPIPS, CLIP-Score, DINOv2) throughout this process—particularly their alignment with human perceptual judgments. Method: Combining controlled user studies, subjective similarity ratings, and rigorous statistical analysis, the work quantifies how progressive manual prompt refinement affects regeneration fidelity. Contribution/Results: It provides the first empirical validation that iterative human-in-the-loop prompt engineering significantly improves image regeneration quality—both subjectively and across most objective metrics. CLIP-Score demonstrates strong correlation with human judgments (r > 0.85), whereas LPIPS shows notably weaker alignment. The study establishes the first empirically grounded benchmark for human–machine similarity alignment, offering both methodological foundations and evaluation standards for interpretable AI image editing and human–AI collaborative prompt engineering.

Assesses alignment of image similarity metrics with human perceptionEvaluates iterative human-driven prompt refinement for AI image regenerationExamines effectiveness of incremental prompt adjustments in improving output

Nepotistically Trained Generative-AI Models Collapse

Nov 20, 2023
MB
Matyáš Boháček
🏛️ Stanford University | University of California, Berkeley

This work identifies a systematic degradation phenomenon—termed “nepotistic training”—occurring when generative AI models are fine-tuned using images synthesized by themselves. Leveraging diffusion-based text-to-image frameworks guided by CLIP, the study conducts controlled retraining experiments, multi-dimensional quality assessments (e.g., FID), and distributional shift analyses. It empirically demonstrates that the degradation is (i) contagious—injecting merely 0.5% AI-generated images degrades FID by over 200%; (ii) generalizable—distortions persist across unseen prompts; and (iii) irrecoverable—subsequent fine-tuning on clean real data fails to restore performance. The paper formally defines and validates these three core properties of this novel training failure mode. By establishing both theoretical insight and empirical evidence, the findings provide critical implications for sustainable generative model training and responsible content governance.

AI models distort images when retrained on their own outputsDistortion affects unrelated text prompts after retrainingModels fail to fully recover even with real data retraining

Latest Papers

What's happening recently
View more

This work addresses the challenge that non-expert users often struggle to articulate precise aesthetic intentions in traditional photo editing, which typically relies on explicit user instructions. To overcome this limitation, we propose SmartPhotoCrafter—the first unified framework integrating reasoning and generation for automatic photographic image enhancement. Our approach employs an Image Critic module to automatically detect visual flaws and a Photographic Artist module to perform targeted refinements, eliminating the need for manual guidance. Trained via a multi-stage strategy—including base pretraining, reasoning-guided supervised fine-tuning, and joint reinforcement learning—our model achieves controllable generation and semantic consistency on a newly curated staged dataset. Experiments demonstrate that SmartPhotoCrafter outperforms existing generative models in automatic photo enhancement, producing results with superior photorealism and heightened sensitivity to tonal adjustments.

aesthetic understandingautomatic enhancementimage quality

Existing generative image editing methods often rely on retraining or fine-tuning, which incurs high computational costs and operational complexity. This work proposes EPEdit, a zero-shot image editing system that requires no fine-tuning and uniquely integrates the zero-shot editing capabilities of Stable Diffusion with an intuitive interactive design, supporting diverse editing tasks guided by text prompts and user-provided masks. Built upon a lightweight client-server architecture, EPEdit maintains computational efficiency while substantially lowering the barrier to entry for end users. User studies demonstrate that EPEdit consistently outperforms existing tools in terms of editing quality, subject consistency, and overall performance, offering both usability and efficiency.

cost-effectivegenerative AIimage editing

Existing text-to-image systems often prematurely commit to fine-grained details, constraining early-stage creative exploration and causing uncontrolled alterations during editing, which undermines users’ sense of agency and creative ownership. To address this, this work proposes Creo, a multi-stage text-to-image generation framework that supports progressive human-AI co-creation. Creo introduces intermediate abstract representations to enable phased refinement—from sketch to high-resolution output—allowing users to lock in confirmed decisions at each stage and apply localized, differential updates to specific regions or attributes without triggering global re-rendering and associated semantic drift. Experimental results demonstrate that Creo significantly enhances users’ perceived ownership and output diversity, outperforming one-shot generation baselines in controllability, creative expressiveness, and result richness.

creative ideationgenerative systemsimage editing

This work addresses the challenge of open-domain image generation, where models must handle diverse and complex user requests yet struggle to generalize effectively or evolve autonomously. To this end, we propose GenEvolve, a novel framework that formulates the generation process as a trajectory of tool orchestration. By comparing multiple trajectories for the same request, GenEvolve extracts structured visual experience and employs a privileged teacher branch to guide dense token-level self-distillation in a student model. Our approach introduces the first unsupervised visual experience distillation mechanism based on trajectory contrast, integrated with a procedural prompt-reference construction paradigm that jointly optimizes reference selection and prompt formulation. Evaluated on established benchmarks and our newly introduced GenEvolve-Bench, GenEvolve substantially outperforms strong baselines, achieving state-of-the-art performance.

agentimage generationself-evolving

Hot Scholars

YW

Yunchao Wei

Professor, Beijing Jiaotong University, UTS, UIUC, NUS
Computer VisionMachine Learning
PT

Philip Torr

Professor, University of Oxford
Department of Engineering
GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
YH

Yuhao Huang

Shenzhen University
Medical Image ComputingUltrasoundModel Robustness
XM

Xin Meng

University of Pittsburgh
AI and medical imaging