Dr. Free: You Don't Need Difficulty Rewards for Self-Evolving Search Agents

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational cost of difficulty rewards and their inability to distinguish multi-hop evidence in the self-evolution of data-independent search agents. To this end, we propose the first self-evolution framework that eliminates difficulty rewards by introducing a novel evidence-necessity-based reward mechanism. Specifically, our approach leverages knowledge graph relation chains to construct explicit multi-hop structures, integrating graph sampling, paragraph alignment, and teacher-forced likelihood to model information gain. This design directly optimizes evidence necessity, replacing conventional difficulty assessment and repeated solver sampling. Experimental results demonstrate that the proposed method outperforms state-of-the-art approaches across seven open-domain question-answering benchmarks, achieving substantial improvements on multi-hop tasks while reducing training time by more than sevenfold.
📝 Abstract
A central limitation of current data-free self-evolution methods for training search agents is their reliance on difficulty-based proposer rewards. These methods reward a proposer for generating questions that challenge a co-evolving solver, using solver difficulty as a proxy for question quality. Yet difficulty alone is insufficient to distinguish questions that require cross-passage evidence from those that are answerable via simpler shortcuts. In addition, measuring difficulty demands repeated solver rollouts for every candidate question, leading to substantial computational costs. In this paper, we introduce \methodname, the first self-evolving search framework that eliminates difficulty-based proposer rewards and directly optimizes for evidence necessity relative to shortcut contexts. Dr. Free samples relational chains from a knowledge graph and pairs them with aligned passages, giving question generation an explicit multi-hop structure. A generated question receives a positive information-gain reward only when the likelihood of the target answer under the complete evidence passages exceeds the maximum likelihood under all evaluated shortcut contexts. Because this signal is computed from teacher-forced likelihoods, it removes the need for pass-rate estimation and reduces proposer training time by over $7\times$. Experiments on seven open-domain QA benchmarks show that Dr. Free outperforms prior data-free search agents and the supervised baseline, with large improvements on multi-hop QA benchmarks.
Problem

Research questions and friction points this paper is trying to address.

self-evolving search agents
difficulty-based rewards
multi-hop question answering
computational cost
shortcut contexts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-evolving search agents
Difficulty-free rewards
Information-gain reward
Multi-hop reasoning
Knowledge graph