Score
Design, build, or evaluate image synthesis systems that condition generation or editing on externally retrieved information. This includes autonomously planning and executing searches, retrieving out‑of‑distribution facts or assets, and merging retrieved results into generation prompts or edit constraints so produced images are grounded in the retrieved knowledge.
Conditional image synthesis with diffusion models suffers from a lack of systematic understanding due to architectural complexity, task heterogeneity, and diverse conditional mechanisms. Method: This paper introduces the first unified taxonomy for conditional diffusion modeling, categorizing approaches by *where* conditioning is injected—either into the denoising network architecture or the sampling process—and formalizes three paradigmatic stages: training, reuse, and specialization. It further classifies six mainstream sampling-time conditioning strategies. Contributions: Based on a structured analysis of over 100 works, the paper establishes a comprehensive knowledge framework and open-sources an authoritative resource repository (GitHub Awesome-Conditional-Diffusion-Models). It identifies persistent bottlenecks—including limited generalization, inefficient inference, and coarse-grained control—and proposes principled directions toward scalable, modular, and fine-grained conditional modeling.
This work addresses the limitation of existing image generation models, which rely on static internal knowledge and struggle to incorporate external or up-to-date information required in real-world scenarios. The authors propose the first trainable search-augmented image generation agent that leverages multi-hop retrieval to acquire both textual knowledge and reference images, enabling knowledge-grounded image synthesis. To support this approach, they introduce dedicated datasets—Gen-Searcher-SFT-10k and Gen-Searcher-RL-6k—and a new evaluation benchmark, KnowGen. They further design a dual-modality reward mechanism for reinforcement learning, combining supervised fine-tuning with the GRPO algorithm. Experimental results demonstrate substantial improvements, with performance gains of approximately 16 points on KnowGen and 15 points on the WISE benchmark, significantly outperforming strong baselines such as Qwen-Image.
Current image generation models exhibit limitations in handling ambiguous intents, logical reasoning, and out-of-distribution knowledge, primarily due to the absence of deep reasoning capabilities and real-time access to external information. This work proposes the first training-free, plug-and-play multi-stage agent framework that seamlessly integrates retrieval-augmented generation with deep reasoning. By employing a closed-loop “question-answering” mechanism, the framework autonomously identifies gaps in logic and knowledge, dynamically plans retrieval and reasoning steps, and orchestrates them in an adaptive manner. Introducing the agent paradigm into image generation for the first time, the method achieves significant performance gains on the WISE Verified and RISEBench benchmarks, delivering absolute improvements of 0.313 and 19.70 for Qwen-Image and Qwen-Image-Edit-2511, respectively, thereby establishing a new state of the art among open-source models.
Current visual generative models exhibit significant limitations in spatial reasoning, state persistence, long-term consistency, and causal understanding, hindering their ability to produce structurally coherent and intelligently behaving content. This work proposes a paradigm shift from appearance-based synthesis toward intelligent visual generation, introducing a novel five-level generative capability taxonomy—from atomic generation to world modeling—that emphasizes the integration of structure, dynamics, domain knowledge, and causality. By leveraging key technical components including a unified understanding-generation architecture, flow matching, enhanced representations, post-training optimization, and synthetic data distillation, the study establishes a capability-centered evaluation framework. This framework exposes the prevailing overreliance on perceptual quality metrics while neglecting structural and causal deficiencies, thereby charting a roadmap for the development of next-generation intelligent visual generation systems.
This study investigates whether humans can accurately regenerate target images via iterative prompt optimization in AI-based image generation, and systematically evaluates the reliability of mainstream image similarity metrics (LPIPS, CLIP-Score, DINOv2) throughout this process—particularly their alignment with human perceptual judgments. Method: Combining controlled user studies, subjective similarity ratings, and rigorous statistical analysis, the work quantifies how progressive manual prompt refinement affects regeneration fidelity. Contribution/Results: It provides the first empirical validation that iterative human-in-the-loop prompt engineering significantly improves image regeneration quality—both subjectively and across most objective metrics. CLIP-Score demonstrates strong correlation with human judgments (r > 0.85), whereas LPIPS shows notably weaker alignment. The study establishes the first empirically grounded benchmark for human–machine similarity alignment, offering both methodological foundations and evaluation standards for interpretable AI image editing and human–AI collaborative prompt engineering.
This work identifies a systematic degradation phenomenon—termed “nepotistic training”—occurring when generative AI models are fine-tuned using images synthesized by themselves. Leveraging diffusion-based text-to-image frameworks guided by CLIP, the study conducts controlled retraining experiments, multi-dimensional quality assessments (e.g., FID), and distributional shift analyses. It empirically demonstrates that the degradation is (i) contagious—injecting merely 0.5% AI-generated images degrades FID by over 200%; (ii) generalizable—distortions persist across unseen prompts; and (iii) irrecoverable—subsequent fine-tuning on clean real data fails to restore performance. The paper formally defines and validates these three core properties of this novel training failure mode. By establishing both theoretical insight and empirical evidence, the findings provide critical implications for sustainable generative model training and responsible content governance.
This work addresses the challenge that non-expert users often struggle to articulate precise aesthetic intentions in traditional photo editing, which typically relies on explicit user instructions. To overcome this limitation, we propose SmartPhotoCrafter—the first unified framework integrating reasoning and generation for automatic photographic image enhancement. Our approach employs an Image Critic module to automatically detect visual flaws and a Photographic Artist module to perform targeted refinements, eliminating the need for manual guidance. Trained via a multi-stage strategy—including base pretraining, reasoning-guided supervised fine-tuning, and joint reinforcement learning—our model achieves controllable generation and semantic consistency on a newly curated staged dataset. Experiments demonstrate that SmartPhotoCrafter outperforms existing generative models in automatic photo enhancement, producing results with superior photorealism and heightened sensitivity to tonal adjustments.
Existing generative image editing methods often rely on retraining or fine-tuning, which incurs high computational costs and operational complexity. This work proposes EPEdit, a zero-shot image editing system that requires no fine-tuning and uniquely integrates the zero-shot editing capabilities of Stable Diffusion with an intuitive interactive design, supporting diverse editing tasks guided by text prompts and user-provided masks. Built upon a lightweight client-server architecture, EPEdit maintains computational efficiency while substantially lowering the barrier to entry for end users. User studies demonstrate that EPEdit consistently outperforms existing tools in terms of editing quality, subject consistency, and overall performance, offering both usability and efficiency.
本文提出了一种评估视觉生成系统自主性的框架,通过定义控制器在生成过程中的控制层级来解决现有研究中缺乏一致标准的问题。
Existing text-to-image systems often prematurely commit to fine-grained details, constraining early-stage creative exploration and causing uncontrolled alterations during editing, which undermines users’ sense of agency and creative ownership. To address this, this work proposes Creo, a multi-stage text-to-image generation framework that supports progressive human-AI co-creation. Creo introduces intermediate abstract representations to enable phased refinement—from sketch to high-resolution output—allowing users to lock in confirmed decisions at each stage and apply localized, differential updates to specific regions or attributes without triggering global re-rendering and associated semantic drift. Experimental results demonstrate that Creo significantly enhances users’ perceived ownership and output diversity, outperforming one-shot generation baselines in controllability, creative expressiveness, and result richness.
This work addresses the challenge of open-domain image generation, where models must handle diverse and complex user requests yet struggle to generalize effectively or evolve autonomously. To this end, we propose GenEvolve, a novel framework that formulates the generation process as a trajectory of tool orchestration. By comparing multiple trajectories for the same request, GenEvolve extracts structured visual experience and employs a privileged teacher branch to guide dense token-level self-distillation in a student model. Our approach introduces the first unsupervised visual experience distillation mechanism based on trajectory contrast, integrated with a procedural prompt-reference construction paradigm that jointly optimizes reference selection and prompt formulation. Evaluated on established benchmarks and our newly introduced GenEvolve-Bench, GenEvolve substantially outperforms strong baselines, achieving state-of-the-art performance.