agentic image generation

Design, build, and evaluate image-generation agents that plan and execute multi-stage workflows by composing model calls, tools, and execution actions to progressively construct, condition, and edit images without additional training. Implement and analyze the planner, grounding/context-acquisition, and executor modules so the system can coordinate plug‑and‑play steps, handle underspecified or up‑to‑date requests, and produce or modify images based on acquired context.

agenticimagegeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.27
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A Practical Guide for Designing, Developing, and Deploying Production-Grade Agentic AI Workflows

Dec 09, 2025
EB
Eranga Bandara
🏛️ Old Dominion University | Deloitte & Touche LLP | Florida International University | AnaletIQ | IcicleLabs.AI | Nanyang Technological University | University of Colombo | Effectz.AI

Production-grade autonomous AI workflows face significant engineering challenges in reliability, observability, maintainability, and security governance. Method: We propose a structured, full-lifecycle methodology comprising a multi-agent architecture with collaborative reasoning, tool augmentation, and dynamic orchestration—integrated with the Model Context Protocol (MCP), deterministic orchestration, pure function invocation, containerized deployment, and modular tool integration. We further define nine core engineering practices, including tool-first design, single-responsibility agents, externalized prompt management, and model-federation-driven responsible AI design. Contribution/Results: This work establishes the first systematic engineering paradigm for Agentic AI productionization, markedly improving system simplicity, observability, and governability. Empirical validation via a multimodal news analysis–media generation use case demonstrates robustness and scalability. The methodology provides a reusable framework and practical benchmark for industrial-scale autonomous AI systems.

Designing reliable production-grade agentic AI workflowsEnsuring safety, observability, and maintainability in deploymentIntegrating multiple specialized agents with tools and orchestration

Flow: A Modular Approach to Automated Agentic Workflow Generation

Jan 14, 2025
BN
Boye Niu
🏛️ University of Sydney | University of Adelaide | Mohamed bin Zayed University of Artificial Intelligence | Carnegie Mellon University

To address the weak dynamic workflow adaptability and poor fault tolerance of multi-agent systems in complex tasks, this paper proposes an evolvable workflow modeling and runtime re-planning framework based on Activity-on-Vertex (AOV) graphs. We pioneer the representation of workflows as dynamic AOV graphs, integrating dependency analysis with parallelism quantification to assess task dependency complexity and enable modular decoupling. Coupled with large language model (LLM)-based agents, the framework supports real-time rescheduling and autonomous error recovery guided by historical execution performance. Experimental results demonstrate significant improvements: +32.7% in task execution efficiency, +28.4% in goal achievement rate, and enhanced fault tolerance—enabling adaptive adjustment of highly concurrent subtasks under dynamic conditions.

Dynamic AdjustmentMulti-Agent SystemsTask Execution Efficiency

This work addresses the challenge that existing image editing models struggle with abstract, multi-step natural language instructions—such as “make the advertisement more vegan-friendly.” The authors propose an end-to-end learnable long-horizon editing framework that decomposes complex instructions into structured atomic tasks via a planner, while a coordinator dynamically selects appropriate editing tools and target regions. A vision-language critic provides outcome-oriented reward signals to guide the process. Notably, this approach tightly couples task planning with reward-driven execution and iteratively refines the planner using successful trajectories, eliminating reliance on handcrafted rules or behavioral cloning from expert demonstrations. Experiments demonstrate that the framework produces significantly more coherent and reliable edits under complex, open-ended instructions compared to single-step models and rule-based multi-step baselines.

abstract editing tasksimage editinginstruction following

AnimAgents: Coordinating Multi-Stage Animation Pre-Production with Human-Multi-Agent Collaboration

Nov 21, 2025
WW
Wen-Fan Wang
🏛️ National Taiwan University | Google DeepMind | HCI Research | UCLA

In animation pre-production, fragmented generative AI tools hinder cross-stage collaboration (conceptualization, scripting, design, storyboarding), resulting in content inconsistency and weak creative control. This paper proposes a multi-agent collaborative system tailored for animation pre-production, centered on a coordinating agent that orchestrates stage-specific agents. The system integrates stage-aware task scheduling with an element-level editable, visual dashboard to enable unified workflow orchestration and human-AI co-control. By synergizing generative AI, dynamic task orchestration, and structured information management, it significantly enhances cross-stage coherence and creative controllability. In a controlled study with 16 professional animators, the system outperformed a single-agent baseline across all metrics—collaborative coordination, narrative consistency, information management, and user satisfaction—at *p* < .01. Real-world deployment further validates its practicality and scalability.

Coordinating fragmented outputs across multiple animation stagesIntegrating isolated AI tools into unified pre-production workflowManaging large information volumes while maintaining creative continuity

Existing workflow-based image generation agents struggle to reuse past experiences and user preferences, resulting in low efficiency and poor reliability for repetitive tasks. This work proposes a self-evolving skill mechanism that formulates workflow construction as a typed graph editing task. By integrating staged tool invocation, an automatic rollback mechanism, and a region-level vision-language model (VLM) verifier, the approach translates visual failures into actionable repair suggestions. Structured, reusable skills are continuously refined through joint distillation of execution trajectories, error logs, and VLM feedback. Implemented within the ComfyUI platform, the method achieves state-of-the-art generation scores across four benchmarks, three agent variants, and two image backbones, with human evaluators showing a significant preference for the skill-evolving version.

agent memoryreusable skillsskill evolution

Latest Papers

What's happening recently
View more

This work addresses the limitations of current text-to-image models in open-world scenarios, where they struggle to jointly handle complex semantic understanding, multi-step reasoning, and integration of external knowledge, while lacking a unified agent to coordinate reasoning, tool invocation, and generation. The authors propose a Unified Multimodal Model (UMM) post-training framework that, for the first time, encapsulates the entire image generation pipeline under a single agent policy. They introduce a Reason-Act-Draw GRPO reinforcement learning algorithm, supported by dedicated training infrastructure, combining supervised fine-tuning with reinforcement learning to orchestrate retrieval and generative tools. Their approach incorporates trajectory format conversion and a joint intent–quality reward mechanism. Experiments demonstrate substantial improvements over fixed-pipeline or partial-agent baselines, and the authors release both the training data and the complete post-training framework.

agentic controlmultimodal reasoningopen-world tasks

Existing agents struggle to coordinate diverse visual tools for complex image creation and editing due to the absence of large-scale, executable trajectory supervision. To address this, this work introduces CanvasAgent, an intelligent agent, along with CanvasCraft—the first large-scale dataset of executable trajectories tailored for complex image generation. The approach integrates supervised fine-tuning (SFT) and generalized reinforcement policy optimization (GRPO), incorporating multimodal understanding, intermediate result validation, visual asset tracking, and heterogeneous tool scheduling. A hybrid reinforcement learning strategy is further designed, combining outcome-based and process-based rewards. Experimental results demonstrate that the proposed method significantly outperforms baseline approaches in both final image quality and trajectory plausibility, validating its effectiveness in multi-tool collaborative visual creation.

complex editingexecutable trajectoriesimage creation

Existing vision-based tool-using agents lack a unified, realistic, and state-aware evaluation benchmark, making it difficult to assess their ability to integrate visual understanding with tool execution in multi-image, multi-turn interactions. This work proposes the first unified evaluation framework for vision-guided tool use, featuring a stateful execution environment spanning 16 domains and supporting over 500 tools. Diverse visually grounded tasks are generated through an information-flow-guided automated scene synthesis pipeline coupled with a multi-stage filtering mechanism. Benchmarking 12 leading models reveals that even the strongest model achieves a success rate below 50%, with 53% of failures attributable to errors in image information extraction—indicating that the current performance bottleneck has shifted from planning capability to visual perception accuracy.

information extractionmulti-turn tasksvisual precision

Hot Scholars

CF

Chelsea Finn

Stanford University, Physical Intelligence
machine learningroboticsreinforcement learning
KL

Kangwook Lee

University of Wisconsin-Madison, KRAFTON AI
Machine LearningInformation Theory
ZZ

Zhaoxiang Zhang

Institute of Automation, Chinese Academy of Sciences
Computer VisionPattern RecognitionBiologically-inspired Learning
LH

Lifang He

Associate Professor of Computer Science, Lehigh University
Machine LearningAI for HealthMedical ImagingBiomedical Informatics