ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL

📅 2026-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决开放性任务中强化学习奖励难以获取的问题,提出ArenaFlow框架,通过基于锦标赛的相对排名和分层信用传播来优化行为并提炼可复用技能。
📝 Abstract
Reinforcement learning has substantially improved large language model (LLM) agents in verifiable domains, but remains difficult to apply to open-ended agent tasks, where solutions are diverse and reliable scalar rewards are hard to obtain. Recent pairwise evaluation methods alleviate reward discrimination collapse by replacing pointwise scoring with relative preferences. However, they still compress rich comparative feedback into a single trajectory-level reward, obscuring decisive intermediate steps and preventing successful behaviors from being consolidated into reusable skills. We propose ArenaFlow, a hierarchical credit propagation framework for open-ended agent reinforcement learning. ArenaFlow leverages tournament-based relative ranking to derive trajectory-level reward signals. Each comparison is further equipped with structured reflective evaluation, which reveals three types of supervision: pivotal success steps, reusable strategy skills, and usage attribution of retrieved skills. At the step level, ArenaFlow propagates trajectory-level advantages to high-confidence pivotal steps according to tournament survival depth, enabling more targeted optimization of local reasoning behaviors. At the skill level, ArenaFlow estimates skill utility from group-level usage attribution and maintains a global skill memory through utility-aware updating, pruning, and retrieval. The resulting high-utility skills further serve as policy priors for future exploration. Extensive experiments validate ArenaFlow's effectiveness on open-ended agent tasks.
Problem

Research questions and friction points this paper is trying to address.

reinforcement learning
open-ended agent tasks
trajectory ranking
credit propagation
relative preferences
Innovation

Methods, ideas, or system contributions that make the work stand out.

hierarchical credit propagation
tournament-based relative ranking
structured reflective evaluation
trajectory-level advantage propagation
global skill memory
🔎 Similar Papers
No similar papers found.
Q
Qiang Zhang
Alibaba Token Hub, Alibaba Group
R
Ruixue Ding
Alibaba Token Hub, Alibaba Group
F
Fanrui Zhang
Alibaba Token Hub, Alibaba Group
X
Xi Chen
Alibaba Token Hub, Alibaba Group
Boli Chen
Boli Chen
University College London
Systems and ControlOptimizationSmart Cities
Shihang Wang
Shihang Wang
DAMO Academy, Alibaba Inc.
Natural Language Processing
Y
Yinfeng Huang
Amap, Alibaba Group
Y
Yi Zheng
Amap, Alibaba Group
Pengjun Xie
Pengjun Xie
Alibaba Group
NLP/IR/ML
Kaipeng Zhang
Kaipeng Zhang
Shanghai AI Laboratory
LLMMultimodal LLMsAIGC
J
Jiawei Liu
Alibaba Token Hub, Alibaba Group
Z
Zheng-Jun Zha
Alibaba Token Hub, Alibaba Group