Cache the Encoder Within:Compact, Reusable Memory across LLM Queries

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the redundant encoding and excessive persistent cache storage costs arising from shared documents across multiple queries. We propose EncBank, an architecture that repurposes the lower layers of large language models as reusable encoders, enabling cross-precision sharing via self-distilled suffix adapters without quantization-aware retraining. Furthermore, a CoMem interface is designed to decouple low-level encoding from high-level reading, combined with a 4-bit quantization strategy to compress intermediate states. Evaluations on the Qwen series demonstrate that 4-bit quantization incurs less than one point of performance degradation while reducing GPU memory consumption to 28.1% and accelerating prefill by 1.4×. Additionally, the proposed method successfully passes the majority of tasks in Terminal-Bench, confirming its practical effectiveness for efficient multi-query serving.
📝 Abstract
Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs. Building on CoMem's intermediate-state interface, EncBank treats a pretrained LLM's lower layers as a reusable document encoder and compactly stores their outputs for an adapted upper-layer reader. A self-distilled suffix adapter is shared across storage precisions within each backbone, without quantization-specific retraining. Across five benchmark suites on three Qwen backbones spanning different sizes and full-attention and hybrid architectures, 4-bit storage keeps each reported benchmark aggregate within one score point of native-precision EncBank. In a fixed Qwen3-8B workload, it retains 28.1% of the native-precision persistent GPU store. Separate native-precision controls yield a 1.40x selected-pack prefill speedup over same-evidence, same-adapter text replay, at a 3.12-point RULER accuracy cost. A native-precision Qwen3.8-27B configuration also passes 70 of 89 Terminal-Bench 2.1 tasks. EncBank thus combines reusable computation with compact memory, while task fidelity and end-to-end benefits remain dependent on the workload, preparation costs, and reuse frequency.
Problem

Research questions and friction points this paper is trying to address.

redundant encoding
persistent storage costs
repeated queries
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

reusable encoder
self-distilled suffix adapter
compact memory caching
quantization-aware storage
LLM inference acceleration
🔎 Similar Papers
No similar papers found.