Fusion is the New Mutation: Bandit-Guided Evolution on Workflow Graphs

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of high evaluation costs and parent selection difficulties in agentic workflow evolution by proposing the DAGO framework. DAGO pioneers formulating multi-parent fusion as a contextual bandit arm selection problem, leveraging pretrained embeddings and a Diagonal LinUCB policy to balance exploration and exploitation for identifying optimal parent combinations under limited budgets. It further employs large language models to generate child workflows, utilizing directed acyclic graphs for topological management and efficient evolutionary search. Experimental results demonstrate that DAGO achieves the highest macro-average score across six benchmarks, outperforming AFlow while reducing search overhead by 11.2%, thereby validating the effectiveness of its feedback-driven selection mechanism.
📝 Abstract
Automated agentic workflow optimization relies on costly evaluations, making it essential to allocate a limited evaluation budget effectively. Multi-parent fusion can reuse designs from previously discovered workflows, but identifying promising parent combinations requires learning from limited fusion feedback. We introduce DAGO (Directed Acyclic Graph Optimization), a contextual-bandit-guided framework that learns which parent workflows to fuse under a limited evaluation budget. DAGO formulates each candidate parent combination as an arm, represented by pretrained embeddings of its constituent workflows'code and prompts. A diagonal LinUCB policy learns a shared reward model across arms and balances exploitation of arms with high predicted offspring quality against uncertainty-driven exploration. After an arm is selected, an LLM generates a child workflow through summary-guided fusion, and the child's validation score serves as the reward for updating the bandit. A shared directed acyclic graph maintains discovered workflows and their multi-parent lineage, providing an expanding pool of parents for subsequent arm proposals. Across six benchmarks covering mathematical reasoning, code generation, and question answering, DAGO achieves the highest macro-average score among the evaluated baselines. Under matched validation-evaluation budgets, it improves over AFlow from 80.3 to 81.7 while reducing aggregate search expenditure by 11.2%. Ablation studies show that LinUCB-guided arm selection outperforms both random selection and its exploration-free variant, supporting the value of feedback-driven selection and exploration-exploitation balance.
Problem

Research questions and friction points this paper is trying to address.

agentic workflow optimization
evaluation budget
multi-parent fusion
automated optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contextual Bandit
Workflow Graph Optimization
Multi-parent Fusion
LinUCB
Directed Acyclic Graph
🔎 Similar Papers
No similar papers found.
Zhiwei Shang
Zhiwei Shang
The Chinese University of Hong Kong, Shenzhen
Robot LearningReinforcement Learning
Jiahang Sun
Jiahang Sun
The Chinese University of Hong Kong, Shenzhen
Large Language ModelsMulti-Armed Bandits
M
Mingrong Gong
College of Computing and Data Science, Nanyang Technological University, Singapore
M
Mingze Kong
School of Data Science, The Chinese University of Hong Kong, Shenzhen, China
Z
Zikun Qu
School of Data Science, The Chinese University of Hong Kong, Shenzhen, China
P
Pingchen Lu
School of Data Science, The Chinese University of Hong Kong, Shenzhen, China
Junhao Dong
Junhao Dong
Nanyang Technological University
AI SafetyRobust AI
Z
Zhipiao Liu
Huawei Technologies Co., Ltd., China
H
Hongwei Yang
Huawei Technologies Co., Ltd., China
G
Guoqing Xie
Huawei Technologies Co., Ltd., China
Y
Yao Shu
Artificial Intelligence Thrust, The Hong Kong University of Science and Technology (Guangzhou), China
Zhongxiang Dai
Zhongxiang Dai
Assistant Professor, The Chinese University of Hong Kong, Shenzhen
Machine LearningData-Centric AILarge Language ModelsMulti-Armed BanditsBayesian Optimization