Affix Cache for Diffusion Large Language Models

๐Ÿ“… 2026-06-26
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the inefficiency of diffusion-based large language models, where bidirectional attention invalidates shared context caches and necessitates full recomputation across requests. To overcome this, we propose ACache, a novel cross-request cache reuse mechanism supporting shared segments at arbitrary positionsโ€”prefix, middle, or suffix. By introducing an anchor token selection algorithm to identify critical tokens, ACache enables partial KV state recomputation while preserving prefix and suffix caches, further co-designed with modern inference engines for efficient deployment. Experimental results demonstrate that ACache effectively balances accuracy and computational overhead, reducing recomputation latency by 56.7%, increasing throughput by 1.71ร—, and decreasing peak memory consumption by 45.8%, significantly outperforming existing approaches.
๐Ÿ“ Abstract
Diffusion Large Language Models (DLLMs) enable non-autoregressive decoding, but efficient inference support remains immature: unlike autoregressive models, whose requests reuse a shared prefix key-value (KV) cache, DLLMs use bidirectional attention, so a shared context's KV states depend on the tokens still being decoded, leaving directly reused caches stale and full recomputation necessary. We present ACache, a cross-request cache reuse mechanism for shared spans, or affixes, at any position: prefix, infix, or suffix. ACache measures the influence of affix tokens on the masked generation region to identify a small request-specific subset as Anchor Tokens, and recomputes only their KV states while reusing the remaining affix cache. Built on state-of-the-art intra-request caching mechanisms, ACache recovers most of the accuracy lost to direct affix-cache reuse on average when recomputing around 20% of affix tokens, and at that budget preserves more accuracy than selection criteria adapted from prior cross-request cache-reuse systems. We co-design ACache with a modern inference engine, whose attention reads each request's recomputed Anchor KV states alongside one affix cache shared across concurrent requests. Against the same system with only intra-request caching, ACache cuts recompute latency by up to 56.7%, translating to as much as 1.71$\times$ end-to-end throughput, while reducing peak KV cache memory by up to 45.8%.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Large Language Models
KV cache reuse
non-autoregressive decoding
inference efficiency
bidirectional attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Large Language Models
Cross-request Cache Reuse
Anchor Tokens
Non-autoregressive Decoding
Inference Optimization
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.