Reinforcement Learning with Verifiable Rewards for Small Search Agents

πŸ“… 2026-09-23
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study investigates the efficacy of Reinforcement Learning with Verifiable Rewards (RLVR) for small-parameter search agents, with a particular focus on reward design. Methodologically, the GRPO algorithm is applied to the Qwen3-0.8B model integrated with an interleaved Wikipedia search tool, systematically evaluating how different reward shapes affect open-domain question answering performance. The findings demonstrate that RLVR can substantially enhance the capabilities of small models without requiring knowledge distillation. Moreover, conventional sparse exact-match rewards prove inadequate for small-model RLVR scenarios. The best experimental run achieves an average exact match of 0.352, representing a 3.8-fold improvement over the baseline. Ultimately, this work establishes principled guidelines for reward design tailored to small-scale language models operating in retrieval-augmented settings.
πŸ“ Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) performs well on problems with clear rewards, such as mathematics and coding, but whether it also works where the reward is less clear remains open. The reason-over-search recipe applies RLVR to open-domain question answering, where retrieval grounds the answer and a match against the reference supplies the reward. So far it has been demonstrated on large models, and below one billion parameters only with distillation from a larger teacher. We test the recipe on a small model. We train Qwen3.5-0.8B with Group Relative Policy Optimization (GRPO) and an interleaved Wikipedia-search tool on MuSiQue, varying only the reward across three shapes over three seeds each, and we evaluate every checkpoint held-out on a seven-benchmark question-answering suite. The recipe works: the best run reaches 0.352 average exact match against a 0.092 untrained floor, a 3.8-fold gain, with no distillation step in the training loop. The reward shape also matters. The Search-R1-faithful exact-match-only reward is the worst of the three at every seed at the matched training horizon, and it is worst even on exact match, the metric it directly optimises. We conclude that the sparse exact-match reward, RLVR's default in mathematics and code, is the wrong starting point for models of this size. The reason-over-search setting can supply a suitable reward for RLVR on small models, but small-model RLVR needs its own reward-design study rather than a scaled-down copy of a large-model recipe.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning with Verifiable Rewards
Small Language Models
Open-domain Question Answering
Reward Design
Reason-over-search
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning with Verifiable Rewards
Small Language Models
Reason-over-Search
Reward Design
Group Relative Policy Optimization
πŸ”Ž Similar Papers
No similar papers found.
G
Gaurisankar Jayadas
Leiden Institute of Advanced Computer Science, Leiden University, Leiden, The Netherlands
Aske Plaat
Aske Plaat
Leiden University
Large Reasoning ModelsMusicReinforcement Learning
Á
Álvaro Serra-Gómez
Leiden Institute of Advanced Computer Science, Leiden University, Leiden, The Netherlands
S
Sandheep P
Leiden Institute of Advanced Computer Science, Leiden University, Leiden, The Netherlands