🤖 AI Summary
This study addresses critical bottlenecks in LLM-based multi-agent systems, including post-hoc evolution, imbalanced credit assignment, and uncalibrated skill admission. To overcome these challenges, this work proposes an online self-evolving graph orchestration paradigm. Methodologically, it rectifies failed steps in real time via execution features and value estimation, and achieves precise credit assignment through an anchor-trajectory-balanced loss function. Furthermore, a verification-based skill admission mechanism is introduced to ensure the reliability of dynamic updates, integrating flow matching regression with sequential hypothesis testing techniques. Experimental results demonstrate that the proposed approach significantly outperforms baseline models across twelve datasets encompassing question answering and mathematical reasoning tasks.
📝 Abstract
In recent years, LLM-based multi-agent systems have been widely applied to orchestrate tool-using agents into executable communication graphs. However, existing self-evolving orchestration still faces key challenges, including post-hoc evolution that revises the team only after the trajectory ends, credit diffusion that gives every action the same terminal advantage under confounded baselines, and skill admission that is uncalibrated and never retired. To address these challenges, we propose EvoSteer, a new paradigm of Online Self-Evolving Graph Orchestration -- the orchestrator builds a running team and repairs its plausible but failing steps from execution features and a learned value estimate. To support this paradigm, we introduce Anchored Trajectory Balance (AnchorTB), a regression-style flow-matching loss that assigns each orchestration action a coefficient by balancing subtrajectories against a frozen reference. Built on the learned flow, we further propose Validated Skill Admission, in which a candidate skill is tried before promotion and promoted only if paired evidence passes a sequential test under a shared nominal testing budget. Moreover, AnchorTB combines measured task-level reference reward statistics with prefix-dependent corrections. Experimental results on twelve datasets show that EvoSteer significantly outperforms baselines across question answering, mathematical reasoning, code generation, and interactive decision making. Our code is available at https://github.com/beita6969/evosteer.