WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient post-training performance gains in tool use caused by isolated scaling environments by proposing WEFT, a framework that jointly optimizes all components of agent interaction systems. Methodologically, it introduces a system-wide coupled evolution mechanism, integrating execution-driven self-evolution with MegaMCP state isolation to construct scalable agent architectures. During training, prefix-preserving sampling and atomic-turn credit assignment are employed to enable stable, large-scale post-training. Experimental results demonstrate that WEFT significantly outperforms baselines of comparable scale on benchmarks such as BFCL V4, achieving improvements of up to 12.27 percentage points for 14B-parameter models.
📝 Abstract
Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in isolation, however, does not guarantee commensurate gains in model performance, because reliable learning signals depend on coherent interactions among all components of the agentic interaction system. To address this problem, we introduce WEFT (Whole-system Evolution For Tool-use Post-training), which couples scalable agentic interaction system construction, execution-driven self-evolution, and stable post-training. WEFT scales agentic interaction system construction across environment breadth, task complexity, and interaction diversity. Execution-driven self-evolution iteratively uses execution traces and state evidence to attribute failures and revise the responsible components, with fresh rollouts evaluating the changes and providing evidence for subsequent evolution rounds. For stable post-training at scale, WEFT addresses both optimization and execution reliability: prefix-preserving sampling retains verified progress and atomic-turn credit assignment localizes learning signals, while MegaMCP maintains isolated, recoverable state across concurrent rollouts over shared tool services. Extensive experiments across various models and benchmarks demonstrate the effectiveness of WEFT for tool-use post-training. WEFT-8B and WEFT-14B outperform all evaluated matched-size environment-scaling baselines on BFCL V4, $τ^2$-Bench, and Claw-Eval. In particular, WEFT-14B improves over Agent-World-14B by 6.41, 2.23, and 12.27 percentage points. WEFT-35B-A3B further extends these gains to more challenging long-horizon workflow benchmarks, including Toolathlon-Verified and AutomationBench.
Problem

Research questions and friction points this paper is trying to address.

tool-use post-training
agentic interaction system
environment scaling
learning signals
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tool-use Post-training
Whole-system Evolution
Execution-driven Self-evolution
Atomic-turn Credit Assignment
MegaMCP
🔎 Similar Papers
2023-08-22Frontiers Comput. Sci.Citations: 866
💼 Related Jobs
No related jobs found.
B
Bo Mao
East China Normal University
Hang He
Hang He
East China Normal University
AI AgentReinforcement LearningVLMIRLLM4SE
L
Linting Wang
Fudan University
L
Lizhi Lin
Shanghai Qiji Zhifeng Co., Ltd
M
Maosen Zhou
Fudan University
G
Guanming Liu
Fudan University
J
Jinxiu Liu
Shanghai Qiji Zhifeng Co., Ltd
Tianyu Huai
Tianyu Huai
East China Normal University
Continual Learning
Chaoyun Zhang
Chaoyun Zhang
Microsoft
GUI AgentLLMCausal InferenceAIOpsSpatio-temporal Modelling
Bingxuan Li
Bingxuan Li
UIUC
K
Kepeng Lei
Shanghai Qiji Zhifeng Co., Ltd
Guanting Dong
Guanting Dong
Remin University of China
LLM Reasoning & AlignmentDeep Search AgentAgentic RL
Z
Zhou Shao
Shanghai Qiji Zhifeng Co., Ltd
R
Rui Zheng
Shanghai Qiji Zhifeng Co., Ltd
H
Hang Yan
Shanghai Qiji Zhifeng Co., Ltd
Jie Zhou
Jie Zhou
East China Normal University
NLPContinuous LearningSentiment AnalysisLLMsInformation Extraction
Chengcheng Wan
Chengcheng Wan
East China Normal University
Software engineeringsystem optimizationmachine learning
T
Tao Gui
Fudan University, Shanghai Innovation Institute
Liang He
Liang He
East China Normal University
Artificial IntelligenceNatural Language ProcessingHuman-in-the-Loop
X
Xipeng Qiu
Fudan University, Shanghai Innovation Institute