π€ AI Summary
This study addresses the computational inefficiency of reasoning models that continue generating tokens after their answers have already converged. We propose a novel answer-stability-based stopping mechanism that learns optimal termination timing from stability signals within complete trajectories. Through constrained optimization, only the end-of-sequence (EOS) token is fine-tuned, strictly preserving the base modelβs predictive distribution without introducing inference overhead or requiring non-standard decoding strategies. This approach accurately identifies termination points while effectively predicting answer correctness, substantially expanding the accuracy-length Pareto frontier. On MATH-500, our method reduces generated tokens by 40% with merely a 0.5% accuracy drop, outperforming an equivalent-data supervised fine-tuning baseline by 6.16 percentage points.
π Abstract
Reasoning models often continue generating after their answers have settled. Settle learns when to stop from answer stability in completed traces. It trains the existing end-of-reasoning token while keeping other predictions close to the base model, and requires only ordinary decoding at inference. On MATH-500 with Qwen3-4B, Settle reduces token count by 40% with a 0.5-percentage-point decrease in accuracy. It gains 6.16 percentage points over supervised fine-tuning on the same traces shortened at their first stable answer, at nearly identical token counts. Its stopping score predicts whether a correct answer will remain correct. Settle extends the accuracy-token-count Pareto frontier of the evaluated stopping methods.