ReSolve: Reusing Candidate Reasoning through Selective Generative Moderation

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational waste and insufficient answer reliability inherent in multi-solution sampling by proposing ReSolve, a training-free framework. Its core innovation lies in introducing a selective generation regulation mechanism that treats candidate reasoning paths as reusable resources. Specifically, model review is triggered only upon path divergence or the absence of valid answers, while the solving process is bounded within an iterative loop to optimize efficiency. Experimental results demonstrate that ReSolve achieves 100% accuracy with zero error reversals on competition-level mathematics tasks. Furthermore, it reduces token consumption by approximately 47% compared to eight-sample self-consistency, while selective regulation decreases token overhead by 54% relative to full regulation. These findings confirm that the proposed framework delivers both highly efficient and reliable reasoning enhancement.
📝 Abstract
Sampling multiple solutions spends computation on intermediate deductions and unfinished arguments as well as final answers. We introduce ReSolve, a training-free inference procedure that reuses this candidate reasoning through selective generative moderation. An answer-distribution controller invokes a model to examine existing derivations when candidates disagree or lack a parseable answer, then incorporates the generated solution into a bounded loop. Under Hybrid scoring on 130 competition-mathematics problems evaluated with two independently sampled candidate pools, ReSolve obtains 100 and 99 correct answers, compared with 91 and 92 for voting over the same four candidates, with no correct-to-incorrect changes relative to that vote in either pool. Eight-sample self-consistency obtains 94 and 96 correct answers while consuming substantially more tokens; ReSolve uses 46.3% and 47.2% fewer tokens in the two evaluations. A controlled ablation removes visible derivations while retaining answer keys, vote counts, and the per-state output-cap rule, reducing accuracy from 100 to 93 correct despite increasing computation. Selective and always-on Uniform moderation both solve 97 problems, while selectivity reduces moderation tokens by approximately 54% and total pipeline tokens by 6.2%. These results support candidate reasoning as reusable inference computation. They do not establish an accuracy advantage over additional sampling or a distinct benefit from specialized route instructions.
Problem

Research questions and friction points this paper is trying to address.

candidate reasoning reuse
computational efficiency
mathematical problem solving
token consumption
sampling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Selective Generative Moderation
Training-free Inference
Candidate Reasoning Reuse
Answer-distribution Controller
Token Efficiency
🔎 Similar Papers
No similar papers found.