Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决LLM推理中自注意力和KV缓存导致的成本增加问题,提出了一种将长上下文压缩成与答案对齐的记忆嵌入的方法,减少推理时间和能耗。
📝 Abstract
Large language model (LLM) inference is constrained by the quadratic scaling of self-attention and the linear scaling of the KV cache, increasing latency, energy consumption, and GPU memory demand as context length scales. Existing soft-compression methods either lack query-guided memory selection at inference time, train without answer-targeted supervision, or couple compression tightly to a specific decoder architecture. We propose a Context-to-Answer-Aligned Memory Compression (CMC) framework, which compresses long input contexts into compact Context Memory Embeddings (CMEs) aligned to any frozen decoder's embedding space, reducing inference costs without modifying decoder weights. CMC introduces a two-tier KV cache that combines question-guided CME selection with a local context window, and trains the compressor with answer-targeted distillation from a frozen LLM. Experiments across nine encoder-decoder combinations and four QA benchmarks show that CMC consistently outperforms the baseline, achieving up to 7.3 EM and 4.0 F1 point gains on SQuAD, while reducing inference time and energy consumption by up to 20% and peak reserved GPU memory by up to 50% at 3,000 generation tokens. Ablation studies confirm that each architectural component and training objective contributes to the performance.
Problem

Research questions and friction points this paper is trying to address.

Large language model
inference
context length
compression
KV cache
Innovation

Methods, ideas, or system contributions that make the work stand out.

Context-to-Answer-Aligned Memory Compression
Context Memory Embeddings
two-tier KV cache
answer-targeted distillation
🔎 Similar Papers
No similar papers found.
M
Md Mostafizer Rahman
Lucy Family Institute for Data & Society, University of Notre Dame, IN, USA
M
Md Faizul Ibne Amin
The University of Aizu, Aizuwakamatsu, Japan
M
Md Shahajada Mia
The University of Aizu, Aizuwakamatsu, Japan
Y
Yutaka Watanobe
The University of Aizu, Aizuwakamatsu, Japan
Fang Liu
Fang Liu
University of Notre Dame
trustworthy machine learningprivacy & synthetic dataBayesian statisticsmissing data analysis