ASCT: Attentive Search over Counterfactual Trees for Credit Assignment in Agentic Reinforcement Learning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the credit assignment challenge in agent reinforcement learning, where terminal utility is difficult to precisely allocate to intermediate steps. To this end, we propose the ASCT framework, which leverages counterfactual tree search to transform multi-step evaluations into local action credits for optimizing PPO policies. Notably, ASCT introduces a novel decoupling mechanism that employs counterfactual search during training while deploying only the actor network during inference. By integrating MCTS and an AgentUCT variant, the framework achieves efficient credit assignment. Experimental results demonstrate that ASCT attains a utility of 0.6187 on the HotpotQA benchmark, significantly outperforming the VinePPO baseline while reducing computational overhead by 50.3%.
📝 Abstract
Terminal utility evaluates a complete agentic workflow, but learning requires credit for the decisions within it. We introduce Attentive Search over Counterfactual Trees (ASCT), a framework that turns training-time multi-step search into local action credit. At actor-visited states, an auxiliary tree evaluates alternative legal actions from the same recoverable prefix. Its action-value table is centered by the frozen actor's probabilities and supplies credit for PPO on actor-sampled trajectories. This protocol connects counterfactual evaluation to policy learning while deploying the actor alone. Uniform, UCT, and cost-aware AgentUCT instantiate the framework. On HotpotQA agentic retrieval-augmented generation, all three improve mean held-out utility over trajectory-return PPO and workflow-adapted VinePPO. Across three seeds, ASCT-AgentUCT reaches 0.6187 utility versus 0.5939 for VinePPO, with gains in answer F1 and execution cost, and uses 50.3% fewer recorded auxiliary Qwen tokens. Transfer and component-description studies examine the learned policies beyond the training setting.
Problem

Research questions and friction points this paper is trying to address.

Credit Assignment
Agentic Reinforcement Learning
Terminal Utility
Multi-step Decision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Counterfactual Trees
Credit Assignment
Agentic Reinforcement Learning
AgentUCT
PPO
🔎 Similar Papers
No similar papers found.
Y
Yang Li
University of Hong Kong (hku.hk)
J
Jinhan Yang
The Chinese University of Hong Kong (cuhk.edu.hk)
H
hai liu
Jiangxi Science and Technology Normal University (jxstnu.edu.cn)
D
Di Wan
X
Xiyu Chen
University of Hong Kong (hku.hk)
Z
Zongsi Xu
University of Hong Kong (hku.hk)
Tuo Zhou
Tuo Zhou
The University of Hong Kong (HKU)
Large Language ModelMulti-Agents System
Sheng Zhong
Sheng Zhong
Nanjing University
computer networkssecurity and privacytheory of computing
S
Sergey Volkov
University of Hong Kong (hku.hk)
Ye Luo
Ye Luo
University of Hong Kong
Machine LearningEconometrics
H
Hao Sun
Shenzhen University (szu.edu.cn)