🤖 AI Summary
This study addresses the challenges of high evaluation costs and parent selection difficulties in agentic workflow evolution by proposing the DAGO framework. DAGO pioneers formulating multi-parent fusion as a contextual bandit arm selection problem, leveraging pretrained embeddings and a Diagonal LinUCB policy to balance exploration and exploitation for identifying optimal parent combinations under limited budgets. It further employs large language models to generate child workflows, utilizing directed acyclic graphs for topological management and efficient evolutionary search. Experimental results demonstrate that DAGO achieves the highest macro-average score across six benchmarks, outperforming AFlow while reducing search overhead by 11.2%, thereby validating the effectiveness of its feedback-driven selection mechanism.
📝 Abstract
Automated agentic workflow optimization relies on costly evaluations, making it essential to allocate a limited evaluation budget effectively. Multi-parent fusion can reuse designs from previously discovered workflows, but identifying promising parent combinations requires learning from limited fusion feedback. We introduce DAGO (Directed Acyclic Graph Optimization), a contextual-bandit-guided framework that learns which parent workflows to fuse under a limited evaluation budget. DAGO formulates each candidate parent combination as an arm, represented by pretrained embeddings of its constituent workflows'code and prompts. A diagonal LinUCB policy learns a shared reward model across arms and balances exploitation of arms with high predicted offspring quality against uncertainty-driven exploration. After an arm is selected, an LLM generates a child workflow through summary-guided fusion, and the child's validation score serves as the reward for updating the bandit. A shared directed acyclic graph maintains discovered workflows and their multi-parent lineage, providing an expanding pool of parents for subsequent arm proposals. Across six benchmarks covering mathematical reasoning, code generation, and question answering, DAGO achieves the highest macro-average score among the evaluated baselines. Under matched validation-evaluation budgets, it improves over AFlow from 80.3 to 81.7 while reducing aggregate search expenditure by 11.2%. Ablation studies show that LinUCB-guided arm selection outperforms both random selection and its exploration-free variant, supporting the value of feedback-driven selection and exploration-exploitation balance.