Score
Design, build, and evaluate image-generation agents that plan and execute multi-stage workflows by composing model calls, tools, and execution actions to progressively construct, condition, and edit images without additional training. Implement and analyze the planner, grounding/context-acquisition, and executor modules so the system can coordinate plug‑and‑play steps, handle underspecified or up‑to‑date requests, and produce or modify images based on acquired context.
Production-grade autonomous AI workflows face significant engineering challenges in reliability, observability, maintainability, and security governance. Method: We propose a structured, full-lifecycle methodology comprising a multi-agent architecture with collaborative reasoning, tool augmentation, and dynamic orchestration—integrated with the Model Context Protocol (MCP), deterministic orchestration, pure function invocation, containerized deployment, and modular tool integration. We further define nine core engineering practices, including tool-first design, single-responsibility agents, externalized prompt management, and model-federation-driven responsible AI design. Contribution/Results: This work establishes the first systematic engineering paradigm for Agentic AI productionization, markedly improving system simplicity, observability, and governability. Empirical validation via a multimodal news analysis–media generation use case demonstrates robustness and scalability. The methodology provides a reusable framework and practical benchmark for industrial-scale autonomous AI systems.
To address the weak dynamic workflow adaptability and poor fault tolerance of multi-agent systems in complex tasks, this paper proposes an evolvable workflow modeling and runtime re-planning framework based on Activity-on-Vertex (AOV) graphs. We pioneer the representation of workflows as dynamic AOV graphs, integrating dependency analysis with parallelism quantification to assess task dependency complexity and enable modular decoupling. Coupled with large language model (LLM)-based agents, the framework supports real-time rescheduling and autonomous error recovery guided by historical execution performance. Experimental results demonstrate significant improvements: +32.7% in task execution efficiency, +28.4% in goal achievement rate, and enhanced fault tolerance—enabling adaptive adjustment of highly concurrent subtasks under dynamic conditions.
This work addresses the challenge that existing image editing models struggle with abstract, multi-step natural language instructions—such as “make the advertisement more vegan-friendly.” The authors propose an end-to-end learnable long-horizon editing framework that decomposes complex instructions into structured atomic tasks via a planner, while a coordinator dynamically selects appropriate editing tools and target regions. A vision-language critic provides outcome-oriented reward signals to guide the process. Notably, this approach tightly couples task planning with reward-driven execution and iteratively refines the planner using successful trajectories, eliminating reliance on handcrafted rules or behavioral cloning from expert demonstrations. Experiments demonstrate that the framework produces significantly more coherent and reliable edits under complex, open-ended instructions compared to single-step models and rule-based multi-step baselines.
In animation pre-production, fragmented generative AI tools hinder cross-stage collaboration (conceptualization, scripting, design, storyboarding), resulting in content inconsistency and weak creative control. This paper proposes a multi-agent collaborative system tailored for animation pre-production, centered on a coordinating agent that orchestrates stage-specific agents. The system integrates stage-aware task scheduling with an element-level editable, visual dashboard to enable unified workflow orchestration and human-AI co-control. By synergizing generative AI, dynamic task orchestration, and structured information management, it significantly enhances cross-stage coherence and creative controllability. In a controlled study with 16 professional animators, the system outperformed a single-agent baseline across all metrics—collaborative coordination, narrative consistency, information management, and user satisfaction—at *p* < .01. Real-world deployment further validates its practicality and scalability.
Existing workflow-based image generation agents struggle to reuse past experiences and user preferences, resulting in low efficiency and poor reliability for repetitive tasks. This work proposes a self-evolving skill mechanism that formulates workflow construction as a typed graph editing task. By integrating staged tool invocation, an automatic rollback mechanism, and a region-level vision-language model (VLM) verifier, the approach translates visual failures into actionable repair suggestions. Structured, reusable skills are continuously refined through joint distillation of execution trajectories, error logs, and VLM feedback. Implemented within the ComfyUI platform, the method achieves state-of-the-art generation scores across four benchmarks, three agent variants, and two image backbones, with human evaluators showing a significant preference for the skill-evolving version.
This work addresses the limitations of current text-to-image models in open-world scenarios, where they struggle to jointly handle complex semantic understanding, multi-step reasoning, and integration of external knowledge, while lacking a unified agent to coordinate reasoning, tool invocation, and generation. The authors propose a Unified Multimodal Model (UMM) post-training framework that, for the first time, encapsulates the entire image generation pipeline under a single agent policy. They introduce a Reason-Act-Draw GRPO reinforcement learning algorithm, supported by dedicated training infrastructure, combining supervised fine-tuning with reinforcement learning to orchestrate retrieval and generative tools. Their approach incorporates trajectory format conversion and a joint intent–quality reward mechanism. Experiments demonstrate substantial improvements over fixed-pipeline or partial-agent baselines, and the authors release both the training data and the complete post-training framework.
Existing agents struggle to coordinate diverse visual tools for complex image creation and editing due to the absence of large-scale, executable trajectory supervision. To address this, this work introduces CanvasAgent, an intelligent agent, along with CanvasCraft—the first large-scale dataset of executable trajectories tailored for complex image generation. The approach integrates supervised fine-tuning (SFT) and generalized reinforcement policy optimization (GRPO), incorporating multimodal understanding, intermediate result validation, visual asset tracking, and heterogeneous tool scheduling. A hybrid reinforcement learning strategy is further designed, combining outcome-based and process-based rewards. Experimental results demonstrate that the proposed method significantly outperforms baseline approaches in both final image quality and trajectory plausibility, validating its effectiveness in multi-tool collaborative visual creation.
Existing vision-based tool-using agents lack a unified, realistic, and state-aware evaluation benchmark, making it difficult to assess their ability to integrate visual understanding with tool execution in multi-image, multi-turn interactions. This work proposes the first unified evaluation framework for vision-guided tool use, featuring a stateful execution environment spanning 16 domains and supporting over 500 tools. Diverse visually grounded tasks are generated through an information-flow-guided automated scene synthesis pipeline coupled with a multi-stage filtering mechanism. Benchmarking 12 leading models reveals that even the strongest model achieves a success rate below 50%, with 53% of failures attributable to errors in image information extraction—indicating that the current performance bottleneck has shifted from planning capability to visual perception accuracy.