Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback

📅 2026-08-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过Write, Execute, Refine框架,利用强化学习从执行反馈中优化技能,解决了代理生成技能效果不佳的问题。
📝 Abstract
Expert-written natural language skills can improve tool-using agents, yet agent-authored skills perform 8-11 points worse than using no skill. This gap suggests that following procedural guidance and improving it from execution evidence are distinct capabilities. Inference time loops can repair skills but do not improve the model that writes the next one. We study how to organize execution experience from intermediate skills into training states for an optimizer. We introduce WER (Write, Execute, and Refine), a multi-phase framework that trains a Skill Optimizer outside a frozen executor. The optimizer proposes skills, a frozen agent executes each repeatedly, and a programmatic verifier scores the outcomes. The scores provide relative credit and select mixed-outcome records. Matched successful and failed trajectories from these records form the next phase's refinement states, so the optimizer learns from the consequences of its earlier outputs. On BFCL v4 multi-turn and tau2-bench, WER improves average Pass@1 over the no-skill baseline by 7.80 and 3.85 points, respectively. Under an identical refinement workflow, it outperforms the same backbone without optimizer training by 9.35 and 10.29 points. The trained 4B optimizer reaches 76.63 percent on BFCL v4, outperforming all evaluated off-the-shelf general-purpose models used as skill optimizers on average.
Problem

Research questions and friction points this paper is trying to address.

Skill Optimization
Reinforcement Learning
Execution Feedback
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Skill Optimization
Execution Feedback
Multi-phase Framework
Programmatic Verifier
🔎 Similar Papers
K
Kang Peng
Harbin Institute of Technology, Shenzhen, China
Z
Zhiwei Zhang
The Chinese University of Hong Kong; MoE Key Laboratory of High Confidence Software Technologies
Y
Yichen Zhang
Harbin Institute of Technology, Harbin, China
Zezhong Wang
Zezhong Wang
Institute of Science Tokyo
VLSI physical design
Y
Yiming Du
The Chinese University of Hong Kong
Geng Tu
Geng Tu
Harbin Institute of Technology (Shenzhen)
NLPdeep learningtext mining
Baojun Wang
Baojun Wang
Huawei Noah’s Ark Lab
NLP
Bin Liang
Bin Liang
The Chinese University of Hong Kong
NLPdata miningmachine learningtext mining
Ruifeng Xu
Ruifeng Xu
Professor, Harbin Institute of Technology at Shenzhen
Natural Language ProcessingAffective ComputingArgumentation MiningLLMsBioinformatics
K
Kam-Fai Wong
The Chinese University of Hong Kong; MoE Key Laboratory of High Confidence Software Technologies