Hierarchical Credit Assignment for RLVR on Fused Gromov-Wasserstein Geometry

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the uniform credit assignment problem in group-based methods for Reinforcement Learning with Verifiable Rewards (RLVR), which arises from neglecting the global semantic novelty of reasoning. To this end, we propose HarA, a hierarchical credit assignment framework that introduces Fused Gromov-Wasserstein (FGW) geometry into RLVR for the first time. By computing the FGW barycenter, HarA precisely quantifies the semantic novelty of reasoning elements and employs an anchor-guided linearization algorithm for efficient optimization. This mechanism dynamically reweights token-level advantages to enhance the exploration capabilities of large language models (LLMs). As a plug-and-play module, HarA significantly outperforms existing methods across multiple benchmarks, effectively improving LLM reasoning performance.
📝 Abstract
Reinforcement learning with verifiable rewards (RLVR) has been shown to improve the reasoning capability of large language models (LLMs) across diverse reasoning tasks. However, group-based RLVR methods, such as GRPO, assign a uniform advantage to all tokens within rollouts of the same outcome. While existing works refine credit assignment of GRPO based on local signals such as token locations or entropy, they often fail to capture the global semantic novelty of a reasoning behavior relative to the current policy. In this work, we propose a hierarchical credit assignment approach for group-based RLVR methods, called HarA, which identifies and encourages semantically novel reasoning behaviors during RLVR. HarA represents each sampled rollout as a distribution over the hidden states and locations of tokens, and computes the Fused Gromov-Wasserstein (FGW) barycenters of all rollouts with the same outcome, capturing the internal reasoning patterns in the latent space under the current policy. The semantic novelty of a reasoning element can then be measured by its contribution to the FGW distance between the current rollout and the barycenter. While solving the FGW formulation is expensive, we introduce an anchor-guided linearization that turns it into a Wasserstein formulation solvable via the Sinkhorn algorithm efficiently. By reweighing token-level advantage of group-based RLVR methods based on the novelty signals, HarA highlights novel reasoning behaviors at flexible granularities to encourage fine-grained LLM exploration. Extensive experiments across three group-based RLVR methods show that our plug-and-play method effectively enhances the exploration of LLMs, outperforming existing methods across diverse reasoning benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning with Verifiable Rewards
Credit Assignment
Semantic Novelty
Large Language Models
Exploration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Credit Assignment
Fused Gromov-Wasserstein
Reinforcement Learning with Verifiable Rewards
Semantic Novelty
Anchor-guided Linearization
🔎 Similar Papers
No similar papers found.