🤖 AI Summary
This study addresses the phase sensitivity introduced by chunked KV cache compression, which causes periodic position-dependent fluctuations in long-context retrieval accuracy that are obscured by averaged metrics. To investigate this, we formally define and quantify phase sensitivity for the first time, reproducing the phenomenon through training Transformers from scratch. By combining causal interventions with gradient dynamics analysis of idealized retrieval models, we elucidate the underlying mechanism by which attention mechanisms specialize in particular phases. Our findings demonstrate that retrieval accuracy can vary by up to 40 percentage points across different phases, revealing asymmetric contributions of attention components to specific phases. Consequently, this work emphasizes that evaluation protocols must span multiple compression phases to mitigate latent positional biases in compressed long-context systems.
📝 Abstract
Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position relative to compression-window boundaries. We uncover a systematic asymmetry in models using such compression: the same information can be easy to retrieve at one phase and difficult at another. We call this periodic variation in retrieval performance phase sensitivity. In large open-weight models with such compression, long-context retrieval accuracy can differ by up to 40 percentage points across phases, revealing periodic weak spots that average benchmark scores can conceal.
To investigate this behavior, we pretrain a family of transformers from scratch across multiple KV-compression designs, reproducing phase sensitivity across the variants. Mechanistic analysis using causal interventions in these models reveals phase specialization: different attention components contribute asymmetrically to retrieving information at different source phases. We further analyze idealized retrieval models, showing how gradient flow dynamics may favor sharp phase specialization. Evaluating models with chunked KV-cache compression thus requires measuring across compression phases: high average accuracy can coexist with systematic positional failures.