F$^{2}$DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决现有奖励模型无法评估DeepSearch全流程复杂性的问题,提出了F2DR框架,从内容、轨迹和答案三个维度进行全面评估。
📝 Abstract
With the widespread industrial deployment of Large Language Models (LLMs), DeepSearch has emerged as the dominant paradigm for resolving complex user queries. It typically operates through an iterative closed-loop workflow consisting of planning and reflection, information retrieval, and answer generation. However, existing reward models (RMs) and evaluation benchmarks are primarily designed for static single-turn tasks, failing to capture the full-pipeline complexity of DeepSearch workflows. To address this limitation, we propose F2DR, a fine-grained full-pipeline DeepSearch reward framework. F2DR evaluates DeepSearch workflows across three dimensions: Content, Trajectory, and Answer, enabling comprehensive process-level assessment. We further construct DeepSearch RM-Bench, a dedicated benchmark for evaluating RMs in DeepSearch scenarios. Extensive experiments demonstrate that F2DR achieves significantly higher evaluation consistency than self-evaluation-based baselines, while DeepSearch RM-Bench exhibits strong discriminative capability across existing open-source RMs. We will publicly release the complete DeepSearch RM-Bench dataset soon.
Problem

Research questions and friction points this paper is trying to address.

DeepSearch
Reward Models
Evaluation Benchmarks
Innovation

Methods, ideas, or system contributions that make the work stand out.

F2DR
Full-Pipeline Reward Framework
DeepSearch Workflows
RM-Bench
B
Bojian Xiong
TJUNLP Lab, Tianjin University, Tianjin, China
W
Wentao Ding
Baidu Inc., Beijing, China
Y
Yujing Lu
Baidu Inc., Beijing, China
S
Shaowei Zhang
TJUNLP Lab, Tianjin University, Tianjin, China
Ling Shi
Ling Shi
Tianjin University
NLPLLM
J
Jing Liao
Baidu Inc., Beijing, China
Y
Yan Wang
Baidu Inc., Beijing, China
Y
Yueyang Zhang
Baidu Inc., Beijing, China
Long Xia
Long Xia
Research Scientist, Baidu
information retrievaldata miningapplied machine learningrecommender system
Z
Zhiyuan Sun
Baidu Inc., Beijing, China
D
Daiting Shi
Baidu Inc., Beijing, China
J
Jingzhou He
Baidu Inc., Beijing, China
Y
Yuqi Ren
TJUNLP Lab, Tianjin University, Tianjin, China
Deyi Xiong
Deyi Xiong
Professor, College of Intelligence and Computing, Tianjin University, China
Natural Language ProcessingLarge Language ModelsAI4Science