🤖 AI Summary
This study addresses the credit assignment challenge in agent reinforcement learning, where terminal utility is difficult to precisely allocate to intermediate steps. To this end, we propose the ASCT framework, which leverages counterfactual tree search to transform multi-step evaluations into local action credits for optimizing PPO policies. Notably, ASCT introduces a novel decoupling mechanism that employs counterfactual search during training while deploying only the actor network during inference. By integrating MCTS and an AgentUCT variant, the framework achieves efficient credit assignment. Experimental results demonstrate that ASCT attains a utility of 0.6187 on the HotpotQA benchmark, significantly outperforming the VinePPO baseline while reducing computational overhead by 50.3%.
📝 Abstract
Terminal utility evaluates a complete agentic workflow, but learning requires credit for the decisions within it. We introduce Attentive Search over Counterfactual Trees (ASCT), a framework that turns training-time multi-step search into local action credit. At actor-visited states, an auxiliary tree evaluates alternative legal actions from the same recoverable prefix. Its action-value table is centered by the frozen actor's probabilities and supplies credit for PPO on actor-sampled trajectories. This protocol connects counterfactual evaluation to policy learning while deploying the actor alone. Uniform, UCT, and cost-aware AgentUCT instantiate the framework. On HotpotQA agentic retrieval-augmented generation, all three improve mean held-out utility over trajectory-return PPO and workflow-adapted VinePPO. Across three seeds, ASCT-AgentUCT reaches 0.6187 utility versus 0.5939 for VinePPO, with gains in answer F1 and execution cost, and uses 50.3% fewer recorded auxiliary Qwen tokens. Transfer and component-description studies examine the learned policies beyond the training setting.