๐ค AI Summary
This study addresses the inefficiency of diffusion-based large language models, where bidirectional attention invalidates shared context caches and necessitates full recomputation across requests. To overcome this, we propose ACache, a novel cross-request cache reuse mechanism supporting shared segments at arbitrary positionsโprefix, middle, or suffix. By introducing an anchor token selection algorithm to identify critical tokens, ACache enables partial KV state recomputation while preserving prefix and suffix caches, further co-designed with modern inference engines for efficient deployment. Experimental results demonstrate that ACache effectively balances accuracy and computational overhead, reducing recomputation latency by 56.7%, increasing throughput by 1.71ร, and decreasing peak memory consumption by 45.8%, significantly outperforming existing approaches.
๐ Abstract
Diffusion Large Language Models (DLLMs) enable non-autoregressive decoding, but efficient inference support remains immature: unlike autoregressive models, whose requests reuse a shared prefix key-value (KV) cache, DLLMs use bidirectional attention, so a shared context's KV states depend on the tokens still being decoded, leaving directly reused caches stale and full recomputation necessary. We present ACache, a cross-request cache reuse mechanism for shared spans, or affixes, at any position: prefix, infix, or suffix. ACache measures the influence of affix tokens on the masked generation region to identify a small request-specific subset as Anchor Tokens, and recomputes only their KV states while reusing the remaining affix cache. Built on state-of-the-art intra-request caching mechanisms, ACache recovers most of the accuracy lost to direct affix-cache reuse on average when recomputing around 20% of affix tokens, and at that budget preserves more accuracy than selection criteria adapted from prior cross-request cache-reuse systems. We co-design ACache with a modern inference engine, whose attention reads each request's recomputed Anchor KV states alongside one affix cache shared across concurrent requests. Against the same system with only intra-request caching, ACache cuts recompute latency by up to 56.7%, translating to as much as 1.71$\times$ end-to-end throughput, while reducing peak KV cache memory by up to 45.8%.