ProRAG: Process-Supervised Reinforcement Learning for Retrieval-Augmented Generation

πŸ“… 2026-01-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the limitations of traditional outcome-based reinforcement learning in retrieval-augmented generation, where sparse rewards and ambiguous credit assignment often lead models to arrive at correct answers through flawed reasoning pathsβ€”a phenomenon known as process hallucination. To mitigate this, the authors propose ProRAG, a novel framework that integrates online policy exploration with learnable step-level process rewards. ProRAG employs a four-stage pipeline: policy warm-up, Monte Carlo Tree Search (MCTS)-based process reward modeling, reward-guided reasoning refinement, and reinforcement training with a dual-granularity advantage mechanism. This approach enables precise feedback on each action within long-horizon reasoning trajectories, effectively decoupling local behaviors from global outcomes. Evaluated across five multi-hop reasoning benchmarks, ProRAG significantly outperforms both outcome-driven and process-aware baselines, demonstrating particularly strong gains on complex, long-chain reasoning tasks.

Technology Category

Machine Learning: Reinforcement LearningMultiagent Systems: Adversarial AgentsSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
πŸ“ Abstract
Reinforcement learning (RL) has become a promising paradigm for optimizing Retrieval-Augmented Generation (RAG) in complex reasoning tasks. However, traditional outcome-based RL approaches often suffer from reward sparsity and inefficient credit assignment, as coarse-grained scalar rewards fail to identify specific erroneous steps within long-horizon trajectories. This ambiguity frequently leads to"process hallucinations", where models reach correct answers through flawed logic or redundant retrieval steps. Although recent process-aware approaches attempt to mitigate this via static preference learning or heuristic reward shaping, they often lack the on-policy exploration capabilities required to decouple step-level credit from global outcomes. To address these challenges, we propose ProRAG, a process-supervised reinforcement learning framework designed to integrate learned step-level supervision into the online optimization loop. Our framework consists of four stages: (1) Supervised Policy Warmup to initialize the model with a structured reasoning format; (2) construction of an MCTS-based Process Reward Model (PRM) to quantify intermediate reasoning quality; (3) PRM-Guided Reasoning Refinement to align the policy with fine-grained process preferences; and (4) Process-Supervised Reinforcement Learning with a dual-granularity advantage mechanism. By aggregating step-level process rewards with global outcome signals, ProRAG provides precise feedback for every action. Extensive experiments on five multi-hop reasoning benchmarks demonstrate that ProRAG achieves superior overall performance compared to strong outcome-based and process-aware RL baselines, particularly on complex long-horizon tasks, validating the effectiveness of fine-grained process supervision. The code and model are available at https://github.com/lilinwz/ProRAG.
Problem

Research questions and friction points this paper is trying to address.

Retrieval-Augmented Generation
Reinforcement Learning
Process Hallucination
Credit Assignment
Reward Sparsity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Process Supervision
Reinforcement Learning
Retrieval-Augmented Generation
Step-level Reward
MCTS-based Reward Model
πŸ”Ž Similar Papers
No similar papers found.
Z
Zhao Wang
Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China
Z
Ziliang Zhao
Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China
Zhicheng Dou
Zhicheng Dou
Renmin University of China
Information RetrievalRetrieval Augmented GenerationLarge Language ModelsGenerative IR