The Model Knows When to Stop: Training-Free Early Stopping for Long-Context Reading

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational inefficiency of continuous reading in long-context processing and the limitations of existing stopping mechanisms that rely on additional training or exhibit insufficient reliability. To this end, it proposes Answer-Convergence Stopping (ACS), a training-free strategy that introduces a novel stopping criterion based on output signal convergence. By monitoring the log-probabilities of tokens generated by frozen large language models alongside answer-state stability, ACS achieves adaptive early stopping using only output texts and probabilities, without requiring auxiliary components, thereby ensuring cross-model generalizability. Experimental results demonstrate that ACS attains accuracy comparable to full-context reading on LongBench-v2, while achieving premature stopping rates as low as 0%–12% on the S-NIAH benchmark, significantly outperforming conventional methods.
📝 Abstract
Language models often process long inputs sequentially in chunks, but continuing to read after sufficient evidence has been acquired wastes computation. Existing stopping mechanisms either learn sufficiency from internal activations or train an exit gate, while a simpler alternative asks the model whether it has read enough. We introduce Answer-Convergence Stopping (ACS), a training-free stopping rule that measures rather than asks. After each chunk, it probes the frozen model's current answer state and stops when that state is both confident and stable. The rule requires only output-side generation and token log probabilities, has no trained components, and uses one shared configuration across models and benchmarks. Because a stopping policy can save computation simply by stopping too early, we evaluate the stopping decision itself using evidence position where available. On the full LongBench-v2 with two frontier models, ACS is the only stopping policy that matches or exceeds full-reading accuracy. Furthermore, across 250 S-NIAH questions, the premature stopping rate for ACS across five models from two families ranges from 0% to 12%, compared to 8.4% to 45.6% for the verbalized gate. Taken together, ACS reveals that by properly utilizing the output signals of frozen models, we can achieve favorable behaviors like adaptive stopping without the need for additional training.
Problem

Research questions and friction points this paper is trying to address.

Long-context reading
Early stopping
Training-free
Language models
Computational efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-free early stopping
Answer-Convergence Stopping
Long-context reading
Log probabilities
Adaptive computation
🔎 Similar Papers
No similar papers found.
M
Muath Alyobi
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
M
Mohamed Eltahir
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
A
Almoayyad Abuljdail
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
R
Riyadh Almutawa
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
Tanveer Hussain
Tanveer Hussain
Lecturer at Department of Computer Science, Edge Hill University
Computer VisionVideo SummarisationSaliency DetectionFire/Smoke Detection
N
Naeemullah Khan
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia