EvoSteer: Online Self-Evolving Graph Orchestration via Reference-Anchored Credit Assignment

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses critical bottlenecks in LLM-based multi-agent systems, including post-hoc evolution, imbalanced credit assignment, and uncalibrated skill admission. To overcome these challenges, this work proposes an online self-evolving graph orchestration paradigm. Methodologically, it rectifies failed steps in real time via execution features and value estimation, and achieves precise credit assignment through an anchor-trajectory-balanced loss function. Furthermore, a verification-based skill admission mechanism is introduced to ensure the reliability of dynamic updates, integrating flow matching regression with sequential hypothesis testing techniques. Experimental results demonstrate that the proposed approach significantly outperforms baseline models across twelve datasets encompassing question answering and mathematical reasoning tasks.
📝 Abstract
In recent years, LLM-based multi-agent systems have been widely applied to orchestrate tool-using agents into executable communication graphs. However, existing self-evolving orchestration still faces key challenges, including post-hoc evolution that revises the team only after the trajectory ends, credit diffusion that gives every action the same terminal advantage under confounded baselines, and skill admission that is uncalibrated and never retired. To address these challenges, we propose EvoSteer, a new paradigm of Online Self-Evolving Graph Orchestration -- the orchestrator builds a running team and repairs its plausible but failing steps from execution features and a learned value estimate. To support this paradigm, we introduce Anchored Trajectory Balance (AnchorTB), a regression-style flow-matching loss that assigns each orchestration action a coefficient by balancing subtrajectories against a frozen reference. Built on the learned flow, we further propose Validated Skill Admission, in which a candidate skill is tried before promotion and promoted only if paired evidence passes a sequential test under a shared nominal testing budget. Moreover, AnchorTB combines measured task-level reference reward statistics with prefix-dependent corrections. Experimental results on twelve datasets show that EvoSteer significantly outperforms baselines across question answering, mathematical reasoning, code generation, and interactive decision making. Our code is available at https://github.com/beita6969/evosteer.
Problem

Research questions and friction points this paper is trying to address.

multi-agent systems
graph orchestration
self-evolving
credit assignment
skill admission
Innovation

Methods, ideas, or system contributions that make the work stand out.

Online Self-Evolving Graph Orchestration
Anchored Trajectory Balance
Credit Assignment
Validated Skill Admission
LLM-based Multi-Agent Systems
🔎 Similar Papers
Mingda Zhang
Mingda Zhang
Google DeepMind
Computer VisionMulti-modal UnderstandingVideo Generation
H
Hanwen Zhang
Dalian University of Technology, China
Q
Qiang Huang
Fudan University, China
Z
Zijia Wang
University of Oxford, UK
P
Pengfei Guo
North China Electric Power University, China
Yuchen Zhang
Yuchen Zhang
King Abdullah University of Science & Technology
wireless communicationssignal processingconvex optimization
J
Jionghao Zhu
The Chinese University of Hong Kong, Shenzhen, China
X
Xiaoying Tang
The Chinese University of Hong Kong, Shenzhen, China