🤖 AI Summary
This study addresses the computational inefficiency of continuous reading in long-context processing and the limitations of existing stopping mechanisms that rely on additional training or exhibit insufficient reliability. To this end, it proposes Answer-Convergence Stopping (ACS), a training-free strategy that introduces a novel stopping criterion based on output signal convergence. By monitoring the log-probabilities of tokens generated by frozen large language models alongside answer-state stability, ACS achieves adaptive early stopping using only output texts and probabilities, without requiring auxiliary components, thereby ensuring cross-model generalizability. Experimental results demonstrate that ACS attains accuracy comparable to full-context reading on LongBench-v2, while achieving premature stopping rates as low as 0%–12% on the S-NIAH benchmark, significantly outperforming conventional methods.
📝 Abstract
Language models often process long inputs sequentially in chunks, but continuing to read after sufficient evidence has been acquired wastes computation. Existing stopping mechanisms either learn sufficiency from internal activations or train an exit gate, while a simpler alternative asks the model whether it has read enough. We introduce Answer-Convergence Stopping (ACS), a training-free stopping rule that measures rather than asks. After each chunk, it probes the frozen model's current answer state and stops when that state is both confident and stable. The rule requires only output-side generation and token log probabilities, has no trained components, and uses one shared configuration across models and benchmarks. Because a stopping policy can save computation simply by stopping too early, we evaluate the stopping decision itself using evidence position where available. On the full LongBench-v2 with two frontier models, ACS is the only stopping policy that matches or exceeds full-reading accuracy. Furthermore, across 250 S-NIAH questions, the premature stopping rate for ACS across five models from two families ranges from 0% to 12%, compared to 8.4% to 45.6% for the verbalized gate. Taken together, ACS reveals that by properly utilizing the output signals of frozen models, we can achieve favorable behaviors like adaptive stopping without the need for additional training.