Pay for the Fault, Not the Flow: Label-Free In-Flow Multi-Agent Workflow Optimization

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of determining task decomposition granularity, reliance on labeled agent allocation, and high fault-repair costs in multi-agent workflows by proposing the InFlowOp framework. This method introduces a novel label-free inflow optimization mechanism that employs an unsupervised cost function to unify capability-overhead assessment across construction and runtime phases, thereby bidirectionally identifying optimal decomposition granularity and enabling low-cost dynamic error correction during execution. Additionally, this work presents the Braid benchmark for evaluating complex coordination tasks. Experimental results demonstrate that InFlowOp outperforms single-agent baselines by up to 11.97% across multiple domains, with the inflow optimization mechanism contributing a 9.64% performance improvement.
📝 Abstract
Large language models (LLMs) increasingly construct multi-agent workflows that decompose a complex task and assign specialist agents from a pool. However, building such a workflow well remains challenging: how finely to divide the task, which agent to trust with each subtask, and when to create a new specialist are all critical decisions a workflow constructor needs to settle up front. Thus, whether each subtask succeeds remains unknown until the workflow runs. Yet, improving a workflow is costly. Locating a fault usually requires a reference answer, a graded outcome, or a trained assessor, and the fix is applied to the whole workflow through re-execution, re-search, or retraining. We propose InFlowOp, which prices every decision in one label-free cost that weighs how well an agent's competence meets what a subtask demands against how much that agent takes to run. Before execution, InFlowOp bidirectionally determines the granularity of task decomposition and agent assignment following from the cost rather than from a fixed template. During execution, InFlowOp corrects a fault with the cheapest move via the same cost that serves the workflow both as it is built and as it runs. Facing the workflow-level evaluation challenge, we introduce Braid, a benchmark whose tasks require multi-agent coordination beyond single-agent capability. Across various domains and backbones, InFlowOp outperforms single agent baselines by up to $+11.97\%$, achieving $+9.64\%$ with in-flow optimization. Our project page: https://xhguo7.github.io/InFlowOp/.
Problem

Research questions and friction points this paper is trying to address.

multi-agent workflow
workflow optimization
label-free fault detection
task decomposition
benchmark evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-agent workflow optimization
label-free cost estimation
dynamic task decomposition
in-flow fault correction
Braid benchmark
🔎 Similar Papers
No similar papers found.