Questioning the Questions: Sustaining Self-Evolution in Reasoning Models

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation in self-evolving reasoning models caused by invalid problems and diversity collapse. To this end, we propose R-Quest, a framework that guides high-quality self-evolution through validity and novelty feedback. Methodologically, R-Quest introduces a novel mathematical equivalence detection mechanism based on frozen models, combined with solver verification to optimize questioner reward signals and data filtering strategies. Furthermore, it integrates reinforcement learning with semantic similarity computation to achieve efficient data cleansing. Experimental results demonstrate that our approach attains the highest average score across twelve benchmarks, exhibiting consistent performance improvements over ten iterations. Notably, R-Quest outperforms the R-Zero baseline by 17.32 points, effectively mitigating the quality degradation problem inherent in repetitive training cycles.
📝 Abstract
Self-evolving reasoning models learn from their own generated questions, yet repeated self-training can lead to performance collapse. In this paper, we investigate why performance deteriorates over successive rounds and how to sustain self-evolution. Our analysis identifies two recurring quality problems in self-generated questions: invalid questions and repeated variants of the same mathematical questions. First, invalid questions become more prevalent across rounds, and answer-consistency filtering further increases their proportion in training data. Second, existing question diversity controls based on lexical similarity can miss mathematically equivalent questions expressed in different ways, which leads to question diversity collapse in later training rounds. Building on these findings, we introduce R-Quest, which uses question validity and novelty feedback to guide self-evolution. We first train the solver to recognize and reject invalid questions, then use its judgments to guide questioner rewards and filter solver training data. To avoid question repetition, we use a frozen base model to compare sampled question pairs and provide novelty feedback. Empirically, our method consistently achieves the highest average performance on 12 benchmarks in mathematical reasoning, general-domain reasoning, and code generation across two model families. Additionally, R-Quest maintains stable performance gains over ten rounds of self-evolution, peaking in the final round and outperforming R-Zero by 17.32 points.
Problem

Research questions and friction points this paper is trying to address.

self-evolving reasoning models
performance collapse
invalid questions
question diversity collapse
self-training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-evolution
Reasoning models
Question validity
Novelty feedback
Performance collapse