Outcome-based Exploration for LLM Reasoning

📅 2025-09-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Outcome-based reinforcement learning (RL) for enhancing mathematical reasoning in large language models (LLMs) suffers from diversity collapse—over-optimizing answer correctness leads to homogeneous outputs, degrading test-time scaling robustness. Method: We propose a result-space exploration mechanism. First, we identify cross-problem propagation of diversity degradation. Second, we formalize a “result-bandit” theoretical framework and design an Upper Confidence Bound (UCB)-inspired reward that incentivizes exploration over outcome space rather than token sequences. Third, we introduce two complementary algorithms: historical exploration tracking and intra-batch repetition suppression. Results: Experiments on Llama and Qwen demonstrate significant accuracy gains on standard mathematical reasoning benchmarks (e.g., GSM8K, MATH), while effectively mitigating diversity collapse. The approach preserves output heterogeneity and ensures robust generalization under test-time scaling, outperforming prior RL-based fine-tuning methods in both accuracy and diversity metrics.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Search and Optimization: Learning to SearchReasoning under Uncertainty: Stochastic Optimization

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: LLM based quality controls for crowd work
📝 Abstract
Reinforcement learning (RL) has emerged as a powerful method for improving the reasoning abilities of large language models (LLMs). Outcome-based RL, which rewards policies solely for the correctness of the final answer, yields substantial accuracy gains but also induces a systematic loss in generation diversity. This collapse undermines real-world performance, where diversity is critical for test-time scaling. We analyze this phenomenon by viewing RL post-training as a sampling process and show that, strikingly, RL can reduce effective diversity even on the training set relative to the base model. Our study highlights two central findings: (i) a transfer of diversity degradation, where reduced diversity on solved problems propagates to unsolved ones, and (ii) the tractability of the outcome space, since reasoning tasks admit only a limited set of distinct answers. Motivated by these insights, we propose outcome-based exploration, which assigns exploration bonuses according to final outcomes. We introduce two complementary algorithms: historical exploration, which encourages rarely observed answers via UCB-style bonuses, and batch exploration, which penalizes within-batch repetition to promote test-time diversity. Experiments on standard competition math with Llama and Qwen models demonstrate that both methods improve accuracy while mitigating diversity collapse. On the theoretical side, we formalize the benefit of outcome-based exploration through a new model of outcome-based bandits. Together, these contributions chart a practical path toward RL methods that enhance reasoning without sacrificing the diversity essential for scalable deployment.
Problem

Research questions and friction points this paper is trying to address.

Outcome-based RL reduces generation diversity in LLMs
Diversity loss undermines real-world performance and test-time scaling
Reduced diversity on solved problems transfers to unsolved ones
Innovation

Methods, ideas, or system contributions that make the work stand out.

Outcome-based exploration with UCB-style bonuses
Historical exploration for rare answer encouragement
Batch exploration penalizing repetition for diversity