Calibrated e-CUSUM Decoding for Quantized Reasoning Models: Why Token Log-Probability Is the Wrong Observable for Decoding Monitors

📅 2026-07-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the susceptibility of low-bit quantized language models to generation failures caused by degeneration during inference, necessitating effective monitoring mechanisms. The authors propose a training-free decoding controller that constructs a degradation-aware alert score by integrating token-level uncertainty with explicit repetition signals. Leveraging e-process theory, they design a calibrated CUSUM sequential detector for sequence-level monitoring. The study identifies the fundamental inadequacy of centered token log-probabilities as monitoring metrics and instead adopts a historically dependent, autocorrelated, and properly calibrated alternative. Experiments on GSM8K demonstrate that the method improves the accuracy of an INT4 model from 63% to 69% (p=0.18) at a token overhead of 28%, while achieving approximately 60% precision in detecting failing trajectories and substantially mitigating verbatim degeneration.
📝 Abstract
Low-bit quantization makes small reasoning models inexpensive to deploy but can degrade their chains of thought. This motivates decoder-side monitors that intervene when generation becomes unreliable. We show that a natural candidate, the centered token log-probability increment $\log p(w_t)+H_t$, is the wrong observable for this purpose. Under the model's own sampling law it is a mean-zero martingale by construction, so it measures sampling self-consistency rather than trajectory health and is nearly silent during confident repetition, where both $\log p(w_t)$ and entropy are close to zero. We introduce a training-free decoding controller that combines (i) a degeneration-aware alarm score fusing token uncertainty with explicit verbatim repetition and (ii) a calibrated e-process-inspired sequential detector. The raw product process is Ville-valid under a conditional-mean null, while the deployed CUSUM-floored statistic is treated as an empirical change detector because the score is history-dependent and autocorrelated. On GSM8K with DeepSeek-R1-Distill-Qwen-1.5B in FP16 and INT4, calibration turns a monitor that fires on 93--95% of generations into a selective detector of failing traces ($φ\approx 0.3$, precision $\approx 0.6$ against a 0.38 base rate). In this pilot, the controller reduces measured verbatim-degeneration signals and yields a positive but statistically inconclusive INT4 accuracy change from 63% to 69% (paired McNemar $p=0.18$, $n=100$), at a 28% token-budget cost. We also find that non-termination, rather than looping, is the dominant failure mode on GSM8K. The main contribution is methodological: an explanation of why centered token log-probability is inadequate for decoder monitoring and a calibrated, cautiously evaluated replacement.
Problem

Research questions and friction points this paper is trying to address.

quantized reasoning models
decoder monitoring
token log-probability
chain-of-thought degradation
generation reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

calibrated e-CUSUM
quantized reasoning models
decoder monitoring
verbatim degeneration
martingale observables
🔎 Similar Papers
No similar papers found.
E
El Hassane Ettifouri
Novelis Research, Paris, France
A
Ayoub Belfatmi
Novelis Research, Paris, France
M
Mahaman Sanoussi Yahaya Alassan
Novelis Research, Paris, France
W
Walid Dahhane
Novelis Research, Paris, France