🤖 AI Summary
This study addresses the optimization difficulties and flat loss landscapes encountered in KV cache compression under frozen models by thoroughly analyzing the optimization bottlenecks of continuous context compression. Methodologically, it proposes a minimalist Perceiver architecture to enable efficient memory operations, replacing conventional complex designs. Through both theoretical analysis and empirical evaluation, this work reveals the optimization limitations of existing architectures and demonstrates that the proposed minimalist approach significantly outperforms the full Perceiver model and mainstream baselines. Across multiple-choice tasks in diverse domains, including finance and law, the method effectively enhances the utility of long-context compression. Ultimately, this research establishes a novel paradigm for efficient inference in large language models.
📝 Abstract
Many works approach continual learning through the lens of infinite context windows. As an agent puts more observation into context (concretely the KV cache), compacting said context is akin to direct memory manipulation, without affecting the base model's weights. Many works pose KV compaction as an optimization problem: learn a smaller set of KV vectors that matches the behavior of the full KV cache. While this preserves base model behavior, optimizing through a frozen base model results in a highly nontrivial optimization problem with a brittle and flat loss landscape. In this paper, we characterize what makes these optimization problems difficult and demonstrate that a heavily simplified Perceiver-based architecture not only matches performance of a full Perceiver transformer in continuous context compaction, but outperforms baselines on compaction utility. Results are presented on MCQ tasks across Finance, Legal, Gutenberg, and Code.