🤖 AI Summary
This study addresses the cross-request coupling introduced by global error-based stopping mechanisms in HQQ-quantized KV caches during batched inference, where irrelevant requests interfere with target outputs. By analyzing the error propagation pathways inherent in shared stopping logic, this work proposes fixed-iteration and local stopping strategies to decouple intra-batch dependencies. Crucially, it reveals that quantized stopping decisions, rather than quantization itself, constitute the root cause of answer drift. This research establishes a novel perspective emphasizing the necessity of auditing stopping logic beyond merely optimizing quantization group configurations. Experimental results confirm that the proposed approach effectively eliminates companion-induced output shifts, while also demonstrating that natural rebatching may still alter model responses, thereby offering essential guidance for the safe deployment of HQQ in production systems.
📝 Abstract
Language-model systems batch questions for throughput, but unrelated questions should not change a target's answer when its input and numerical execution are fixed. We study compression of the key and value cache, which stores attention representations reused during generation. With request-local groups, Transformers'Half-Quadratic Quantization (HQQ) backend updates compression parameters separately but uses a shared average error to decide when all updates stop. Replacing only the question batched with the target changes four-bit HQQ answers in 170/384 test comparisons across two models. Replaying the other execution's update counts reproduces its complete answer and cache fingerprints in every changed pair, in both directions. Computing the stopping mean in FP32 reduces cache differences but leaves answer changes. Native HQQ also changes confirmed numerical correctness in eight arithmetic pairs. Fixed iterations and request-local stopping remove observed companion dependence under matched controls. Request-local stopping remains sensitive to synthetic padding changes at the tensor level. Fixing the original iteration budget removes this decision path without tuning. Neither repair has an established quality advantage, and natural rebatching still changes answers. Request-independence audits must cover stopping decisions as well as quantization groups.