Score
Designs, implements, and evaluates algorithms and systems that compress and consolidate streaming or cached state online to maintain a compact, retrieval-oriented memory while continuously ingesting new observations. This includes cache compression algorithms, incremental memory compression and online consolidation techniques, and analysis of trade-offs among compression ratio, retrieval accuracy, latency, and update cost.
This work addresses the high latency in real-time stream processing caused by tight coupling between state I/O and the data path, which blocks the CPU on the critical path. To mitigate this, the authors propose Keyed Prefetching, a mechanism that extracts state-access keys from upstream operators and proactively prefetches the required state, thereby overlapping I/O with computation to hide latency. Complementing this, a Timestamp-Aware Caching strategy is introduced to efficiently manage prefetched and historical states in memory. Together, these techniques significantly reduce end-to-end latency for long-running real-time queries while maintaining high throughput and effectively decoupling state access from data processing.
This work addresses the performance bottleneck caused by the substantial KV cache generated during reasoning in large language model (LLM) agents, a challenge exacerbated by the inadequacy of existing compression methods that rely on offline assumptions and fail to generalize to online settings. The study presents the first systematic investigation of online KV cache compression for LLM agents, introducing a framework that leverages low-cost proxy queries—under unknown future query conditions—to enable either immediate or deferred compression of historical cache entries. By integrating Token Eviction (TE) with Attention Matching (AM) and designing proxy query strategies such as boundary-aware selection, repeated prefilling, and delayed generation, the authors demonstrate that deferred compression substantially outperforms immediate compression, with proxy query selection being a critical factor. Experiments on BrowseComp-Plus and WideSearch benchmarks show that TE reduces KV cache usage by up to 80% across model scales while preserving task accuracy and improving throughput, significantly surpassing uncompressed baselines.
This work addresses the challenges of building and maintaining KV caches for unbounded video streams, where existing methods are constrained by limited memory and computational resources and exhibit poor generalization under short-sequence training regimes. The authors propose DSCache, a training-free, decoupled caching mechanism that separates historically accumulated KV caches from on-demand, instantaneously generated ones, further enhanced by position-agnostic encoding to support long-sequence extrapolation. As the first approach specifically targeting KV cache construction in streaming video understanding, DSCache is compatible with existing VideoVLLM frameworks and achieves state-of-the-art performance on the Streaming Video QA benchmark, improving average accuracy by 2.5% over prior methods.
This work addresses the high memory overhead incurred by large-scale sequential recommendation systems when processing long user behavior sequences—a challenge often overlooked by existing approaches that predominantly focus on accuracy while neglecting storage costs. To bridge this gap, we propose MALLOC, the first memory-aware benchmark for comprehensive evaluation of long-sequence compression techniques tailored to recommender systems. MALLOC systematically integrates and adapts compression strategies originally developed for large language models, embedding them into mainstream sequential recommendation architectures. Through reproducible, multi-dimensional assessments encompassing accuracy, efficiency, and computational complexity, MALLOC establishes a standardized framework for evaluating memory-performance trade-offs, thereby filling a critical void in systematic benchmarking and demonstrating its effectiveness in balancing memory efficiency with model performance.
Existing evaluations of external memory systems predominantly rely on static setups, failing to capture the dynamic interplay of streaming memory updates, insertions, and retrievals characteristic of real-world scenarios, thereby yielding misleading performance assessments. This work proposes Neuromem—the first fine-grained, lifecycle-decomposed framework tailored for streaming external memory—which disentangles memory mechanisms into five orthogonal dimensions: data structure, normalization, integration strategy, query construction, and context fusion, and implements a unified benchmarking platform supporting modular component substitution. Experiments across LOCOMO, LONGMEMEVAL, and MEMORYAGENTBENCH reveal that scaling memory size generally degrades performance, time-sensitive queries pose the greatest challenge, the choice of memory data structure primarily dictates performance ceilings, and aggressive compression or generative fusion techniques largely shift costs between insertion and retrieval without substantially improving accuracy.
This work addresses the high computational and storage costs, methodological fragmentation, and lack of a unified theoretical foundation in memory management for large language model agents. Framing multi-level memory compression through rate-distortion theory, the study formulates it as an optimization problem of preserving task-critical information under resource constraints. It introduces a general compression objective and a layer-agnostic lower bound, establishes a seven-dimensional taxonomy, and achieves, for the first time, cross-layer transfer of compression mechanisms. The analysis reveals fundamental limitations of attention magnitude and temporal decay as universal retention signals. Furthermore, the authors develop a unified evaluation benchmark with reference experiments, distill design principles for memory compression in multi-turn interactions, and systematically outline open challenges in the field.
Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.
This work revisits the classical Sleator–Tarjan paging model, which assumes that every first access to newly generated data incurs a compulsory cold page fault. Recognizing that in real systems such data typically resides initially in the processor and thus does not trigger a page fault upon first access, the authors propose a refined model that waives this initial fault cost. Under this more realistic setting, they re-examine online paging algorithms using competitive analysis, offline optimal algorithms (e.g., LFD), and cache-model reformulation. Their analysis demonstrates that no online paging algorithm—whether randomized or augmented with additional resources—can achieve competitiveness in the revised model. This finding suggests that the apparent predictive success of classical competitive analysis stems from an artifact of the original model’s unrealistic assumptions, thereby challenging the theoretical foundations of frameworks such as Cache-Oblivious algorithms.