🤖 AI Summary
This study addresses the inefficiency and parameter generation limitations of conventional tool-based image editing methods that rely on autoregressive models. To overcome these bottlenecks, this work reformulates tool-based editing as a flow matching problem for the first time, proposing a non-autoregressive framework based on conditional rectified flows that directly models high-quality tool parameter distributions. Methodologically, the approach integrates a vision-language model backbone with a diffusion Transformer parameter generator, complemented by a two-stage supervised curriculum and reward-based post-training strategy. Experimental results demonstrate that the proposed framework outperforms specialized multimodal large language model agents across multiple benchmarks while reducing inference latency by 50× and decreasing GPU memory requirements by nearly 2×.
📝 Abstract
Tool-based image editing (image retouching) is commonly formulated with autoregressive multimodal large language models (MLLMs) that sequentially generate reasoning, tool selections, and parameter values. In this work, we present a novel approach to tool-based image editing by framing the task as a flow matching problem. We introduce FlowTool, a framework that directly models the distribution of high-quality tool parameters conditioned on the input image and user instruction using conditional rectified flow. FlowTool combines a vision-language model backbone for multimodal understanding with a Diffusion Transformer parameter generator that transforms Gaussian noise into an editing plan. We train FlowTool with a two-stage supervised flow-matching curriculum, followed by reward-based post-training. Across MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, and MIT-Adobe5K, FlowTool achieves significantly stronger reference-based performance than specialized MLLM editing agents and proprietary MLLMs, while remaining competitive with proprietary models under reference-free evaluation. Moreover, FlowTool significantly improves inference efficiency, reducing latency by at least $50\times$ while requiring nearly $2\times$ less memory than the compared baselines. These results demonstrate that tool-based image editing can be effectively modeled as conditional generation over structured continuous editing parameters, without autoregressive reasoning.