🤖 AI Summary
This study addresses the limitation of existing multi-agent methods that generate only fixed, task-level workflows, rendering them inadequate for handling varying query difficulties and mixed tasks. To overcome this, we propose a training-free dynamic collaboration framework that introduces a novel query-level workflow generation mechanism. By integrating object-oriented agent modeling with execution feedback-driven skill library accumulation, the framework leverages in-context learning to adaptively construct agent ensembles and coordination processes at the individual query granularity, enabling continuous optimization without gradient updates. Experimental results demonstrate that our approach achieves an accuracy of 89.6% on a mixed-task benchmark, outperforming the strongest baseline by 18.1 percentage points. Furthermore, when equipped with a more capable backbone model, the accuracy improves to 92.4%.
📝 Abstract
Multi-agent systems (MAS) powered by large language models have shown strong performance across code generation, mathematical reasoning, and question answering. However, existing methods for automating MAS design mostly operate at the task level, producing a single fixed workflow per benchmark that is applied uniformly to all queries. This assumption fails under realistic conditions. Query difficulty varies widely within a task, and real-world workloads mix heterogeneous task types. We introduce OOPMAS, a training-free framework that generates both the agent set and the coordination workflow at the granularity of individual queries. Agents are represented as object-oriented class definitions with dedicated roles, tools, and persistent state, and workflows are expressed as executable main functions over these agent objects. A dynamic skill library accumulates structured lessons from execution feedback across optimization rounds, enabling in-context improvement without any gradient updates or fine-tuning. On a mixed-task benchmark of queries spanning code, math, and QA, OOPMAS achieves 89.6% accuracy, outperforming the strongest baseline by 18.1 percentage points. A model-swap study across four LLM backbones shows consistent scaling, reaching 92.4% with the strongest model.