Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究解决了LLM代理强化学习中的反馈利用和轨迹聚合问题,通过提出BATON框架,结合贝叶斯反馈归因和轨迹质量标准化方法来优化策略。
📝 Abstract
Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and how complete trajectories are aggregated across a batch. We formulate these dimen- sions as Intra-Trajectory Feedback Attribution and Inter-Trajectory Objec- tive Aggregation, and introduce BATON (Bayesian Attribution and Trajectory Objective Normalization), a dual-axis policy optimization framework. BATON instantiates the first axis with Bayesian Feedback Attribution, which constructs a feedback-conditioned posterior over sampled actions, and the second with Trajec- tory Mass Normalization (TMN), which assigns equal optimization mass to com- plete trajectories. Experiments with GRPO and GiGPO on ALFWorld, WebShop, and SearchQA show that both axes provide independent gains and that their combi- nation consistently achieves the strongest overall performance across model scales.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
LLM Agents
Optimization Dimensions
Feedback Attribution
Trajectory Aggregation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-Axis Policy Optimization
Bayesian Feedback Attribution
Trajectory Mass Normalization
Intra-Trajectory Feedback Attribution
Inter-Trajectory Objective Aggregation
Y
Yingxuan Zhuang
Zhejiang University
B
Binhe Yu
Zhejiang University
J
Jingxiao Yang
Zhejiang University
R
Ruopei Sun
University of Science and Technology of China
Z
Ziting Li
University of New South Wales
C
Cheng Tan
Zhejiang University
Xuhong Zhang
Xuhong Zhang
Zhejiang University
LLMVLMVLATrustworthy AI
Jianwei Yin
Jianwei Yin
Professor of Computer Science and Technology, Zhejiang University
Service ComputingComputer ArchitectureDistributed ComputingAI
J
Jintao Chen
Zhejiang University