RewardWeaver: Long-Horizon Interactive Learning for Language Agents via Self-Evolving Reward Adaptation

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of credit assignment due to sparse terminal feedback and the inability of static process rewards to adapt to evolving policies during long-horizon interactions of language agents. To overcome these bottlenecks, we propose a self-evolving reward adaptation framework that dynamically couples policy optimization, failure attribution, and reward adjustment through a closed-loop mechanism. Specifically, we introduce a semantically anchored verification capability space, employ outcome-driven backward attribution to aggregate common bottlenecks, and incorporate controlled expansion to handle uncovered failure modes, thereby enabling precise process reward allocation. Experiments demonstrate that our method achieves state-of-the-art performance on benchmarks such as SOTOPIA. Furthermore, ablation studies confirm the critical roles of dynamic reward allocation and a stable semantic space in driving overall effectiveness.
📝 Abstract
Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in domains where task outcomes can be reliably evaluated, but long-horizon interaction remains challenging due to sparse terminal feedback and difficult credit assignment. Process rewards provide denser supervision, yet the capabilities most relevant for training can change as the policy evolves: a behavior that is easy to evaluate or frequently deficient need not be the bottleneck currently limiting task success. We introduce RewardWeaver, a self-evolving reward adaptation framework for language agents in long-horizon interaction. RewardWeaver maintains a validated capability space in which the semantics of admitted Rubrics remain fixed, and closes the loop between policy optimization, task evaluation, failure attribution, and reward adaptation. After each training stage, it performs outcome-grounded backward attribution on low-outcome trajectories, aggregates recurrent and policy-controlled capability bottlenecks, and dynamically selects the corresponding process rewards for the next stage. Recurrent failures not covered by the existing capability space trigger a separate, controlled expansion procedure. We evaluate REWARDWEAVER on SOTOPIA, Amazon?HistoryPrice, and a newly constructed Sales Benchmark. Across social interaction, bilateral bargaining, and domain-specific sales, REWARDWEAVER establishes new state-of-the-art (SOTA) results. Ablations further demonstrate the importance of dynamic reward allocation, failure-grounded attribution, and stable semantics for admitted capabilities.
Problem

Research questions and friction points this paper is trying to address.

long-horizon interaction
language agents
sparse terminal feedback
credit assignment
dynamic reward adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Evolving Reward Adaptation
Long-Horizon Interactive Learning
Process Rewards
Failure Attribution
Capability Space
H
Hengbo Xiao
Yoolee.ai
B
Boyao Zhang
University of Science and Technology of China
P
Purui Liu
Peking University
Y
Yuxuan Zheng
Peking University
Haoran Yin
Haoran Yin
Leiden University
H
Haibo Liu
Yoolee.ai
F
Fan Zhang
Yoolee.ai