🤖 AI Summary
This study addresses the reliance of language model reasoning on external verifiers and the problem of erroneous steps contaminating subsequent generation by proposing a self-paced rejection mechanism. The method introduces a novel confidence-bounded first-passage supervision signal, integrating pointwise classification, same-prefix pairwise learning, and group relative policy optimization to train a lightweight LoRA gate that identifies and discards low-recoverability reasoning steps. By freezing the backbone network and employing Monte Carlo dynamic resampling to evaluate prefix recoverability, the approach achieves autonomous error correction without external verification. Across five mathematical benchmarks, it yields average accuracy improvements of 5.4 to 10.1 percentage points with only a 1.21× to 1.40× increase in token overhead, significantly outperforming existing step-level methods.
📝 Abstract
Verifier-guided decoding can prevent harmful reasoning steps from contaminating subsequent generation, but typically relies on an external learned verifier. We ask whether a language model can instead reject its own bad reasoning steps. We define a prefix's recoverability as the probability that the frozen generator can complete it correctly. Diagnostics show that adjacent recoverability changes are often difficult to resolve with practical Monte Carlo budgets, while same-prefix candidates exhibit a sparse low-recoverability tail. We introduce Self-Step Rejection (SSR), which trains a lightweight LoRA acceptance gate on the generator backbone while keeping the base model frozen. SSR uses confidence-qualified first-passage supervision: steps before the first resolved crossing of a root-relative recoverability barrier are accepted, the crossing step is rejected, and unresolved steps and suffixes are excluded. Training combines pointwise classification, same-prefix pairwise learning, and group-relative policy refinement using final-answer correctness. At inference, SSR accepts candidates or resamples from the unchanged prefix under rejection budgets, without an external learned verifier. Across three reasoning models and five mathematical reasoning benchmarks, SSR improves macro-average accuracy over single-pass decoding by 5.4--10.1 points using 1.21--1.40x as many generated tokens, and achieves the highest macro-average accuracy among evaluated step-level methods. Full-solution scaling methods require 4.47--8.27x the single-pass token cost for comparable performance.