GraphHCA: Closed-Form Hindsight Credit Assignment for Long-Horizon LLM Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of step-level credit assignment for long-horizon LLM agents under sparse rewards, where conventional hindsight methods rely on auxiliary models or additional forward passes. We propose a model-free hindsight credit assignment approach that leverages Bayes’ rule to convert hindsight ratios into success probability ratios under the behavior policy. Step-level signals are derived by recursively estimating state potential functions via discounted transitions on a transition graph. This method admits a closed-form solution without learning hindsight distributions, guarantees a unique fixed point on arbitrary directed graphs, and is compatible with GRPO for joint trajectory- and step-level optimization. Empirical evaluations demonstrate state-of-the-art performance on benchmarks such as ALFWorld, improving success rates by 24.6% over GRPO and surpassing the strongest baseline by 4.7%.
📝 Abstract
Group-based reinforcement learning (RL) has advanced large language models (LLMs) and is increasingly extending to agentic tasks, where sparse terminal rewards make step-level credit assignment essential. Existing methods assign credit from what follows an action in sampled rollouts, but do not explicitly capture its retrospective relation to the realized outcome. Hindsight credit assignment (HCA) instead attributes credit through the ratio of hindsight to behavior-policy probabilities, but estimating the hindsight distribution requires an auxiliary model or an extra pass. To address this estimation bottleneck, we propose GraphHCA, a model-free realization of HCA that eliminates explicit hindsight-distribution estimation. For terminal-goal tasks with deterministic transitions, Bayes'rule reduces the hindsight ratio to a ratio of behavior-policy success probabilities at consecutive states. Taking logs yields a state-wise success potential, whose increment across a transition provides step-level credit. GraphHCA estimates this potential from pooled rollouts through a discounted recursion on the induced transition graph, which admits a unique fixed point on any directed graph. The resulting step-level signal is combined with the trajectory-level advantage, requiring neither a learned hindsight model nor an extra forward pass and recovering GRPO when the step-level weight is zero. Among all compared baselines, GraphHCA achieves state-of-the-art results on ALFWorld and WebShop at both LLM scales, and on Sokoban with a vision-language agent. For example, on ALFWorld it improves overall success rate by up to 24.6 points over GRPO and by up to 4.7 points over the strongest step-level baseline.
Problem

Research questions and friction points this paper is trying to address.

Credit Assignment
Hindsight Credit Assignment
Reinforcement Learning
Large Language Model Agents
Sparse Rewards
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hindsight Credit Assignment
Graph-based Recursion
Model-free RL
Step-level Credit
Long-horizon LLM Agents
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.