manage buffer memory

Designs and implements buffer allocation and management systems that provision and resize buffer capacity, orchestrate movement of data between storage and compute, and implement cache eviction and replacement policies. Analyzes local data reuse patterns and access behavior to optimize placement, lifetime, and transfer scheduling of buffered data to reduce latency and memory footprint.

managebuffermemory

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$226K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

The widening gap between CPU and memory access latency poses a fundamental challenge to database performance. Method: This paper systematically surveys four decades of buffer management evolution, analyzing classical algorithms (e.g., LRU-K, 2Q, LIRS, ARC), ML-enhanced approaches, and memory-disaggregated architectures (NVM-aware hierarchies, RDMA-decoupled designs). It introduces the first holistic, time-spanning analytical framework for buffer management evolution and proposes a novel cross-layer adaptive paradigm integrating ML-driven policy learning with eBPF-based kernel extensibility. Contribution/Results: Grounded in empirical analysis of 50+ top-tier conference papers and industrial systems (Linux, PostgreSQL, Oracle), the work distills core design trade-offs and identifies key challenges—including cache coherence across heterogeneous memory tiers, low-overhead decision-making, and OS–DBMS co-design—along with concrete research directions toward scalable, adaptive buffer management.

Analyzing progression from classical algorithms to machine learning policiesIdentifying research challenges in adaptive buffer management for modern systemsSurveying evolution of buffer management algorithms over four decades

Performance Models for a Two-tiered Storage System

Mar 12, 2025
AS
Aparna Sasidharan
🏛️ IIT | Sandia National Lab | Oak Ridge National Lab

To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.

Design and analyze a two-tiered storage systemDevelop online learning for data tier managementEvaluate performance using queuing and behavioral models

Heterogeneous Memory Pool Tuning

May 20, 2025
FV
Filip Vaverka
🏛️ IT4Innovations | VSB - Technical University of Ostrava

To address the challenge of optimizing data placement in HBM+DDR heterogeneous memory systems, this paper proposes a lightweight, non-intrusive, application-level memory analysis and control framework. Leveraging a detailed memory subsystem model on Intel Sapphire Rapids platforms—integrated with empirical bandwidth/latency measurements, runtime allocation tracing, and policy injection—the work quantifies, for the first time, the performance inflection point of HBM-DRAM co-utilization: retaining only 60–75% of critical data in HBM achieves 90% of the platform’s peak performance. The approach requires no source-code modification or recompilation. Evaluated across multiple benchmarks, it significantly improves performance for memory-intensive applications while reducing HBM resource consumption by over 30%. This establishes a deployable, fine-grained data placement optimization paradigm for heterogeneous memory systems.

Analyzing and tuning data placement in heterogeneous memory systemsDetermining optimal data allocation ratios for maximizing platform performanceEvaluating performance of HBM and DDR memory subsystems together

This work addresses the challenges of efficiency, scalability, and consistency in distributed management of KV caches for large language model (LLM) services. It proposes the first four-dimensional taxonomy—spanning locality, lifetime, ownership, and storage substrate—to systematically analyze over 30 existing studies, identifying five architectural paradigms: local paging, decoupled pipelining, shared storage, memory pooling, and hybrid hierarchical designs. The study reveals that “ownership” is a key differentiator in distributed KV cache architectures and highlights the absence of seven KV-specific metrics in current evaluation methodologies. Furthermore, it connects these gaps to six critical open problems, including fault tolerance, isolation, hierarchical eviction, and speculative decoding.

cache managementdistributed memoryKV cache

On Configuring a Hierarchy of Storage Media in the Age of NVM

Apr 16, 2018
SG
Shahram Ghandeharizadeh
🏛️ USC | University of California, Irvine

This work addresses the joint optimization of media selection, capacity allocation, and data placement (replication vs. tiering) for key-value caching across heterogeneous NVM/DRAM/disk storage under memory budget constraints. We introduce the first systematic modeling framework for multi-level non-volatile cache configurations, analytically characterize the operational regimes where replication or tiering dominates, and propose an adaptive configuration policy grounded in device failure rates and data update frequencies. Our methodology integrates cache access behavior modeling, hierarchical configuration optimization, and empirical validation using memcached benchmarks. Results demonstrate that tiering substantially outperforms replication under low device failure rates and high update workloads. Key contributions include: (1) a deployable, low-overhead configuration algorithm; (2) quantitative design guidelines for heterogeneous cache deployment; and (3) theoretical foundations for the reliability–performance trade-off in tiered caching systems.

Determining storage media selection and capacity allocation under budget constraints.Evaluating data replication versus partitioning strategies for performance and recovery.Optimizing memory hierarchy design for caching middleware with NVM and DRAM.

Latest Papers

What's happening recently
View more

Traditional virtual memory–assisted buffer management struggles to efficiently support data migration and access in multi-tier memory architectures. This work proposes vmcacheⁿ, a framework that extends the conventional two-level (DRAM–disk) buffering mechanism to an n-level hierarchy (DRAM–remote memory–disk). Leveraging the operating system’s virtual memory subsystem and page migration facilities, vmcacheⁿ constructs a multi-tier cache pool by integrating remote memory technologies such as NUMA and CXL. To enable fine-grained and low-overhead cross-tier page migration, the framework introduces a new system call, move_pages2. Experimental evaluation under the TPC-C workload demonstrates that vmcacheⁿ achieves up to a 4× improvement in query throughput compared to the original vmcache.

buffer managementn-tierpage migration

Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.

distributed systemsparallel systemsruntime control

Existing burst buffer file systems suffer from performance degradation due to rigid, fixed data layouts that fail to adapt to diverse application I/O behaviors. This work proposes Proteus, a system that treats data layout as a first-class optimization target by integrating static code analysis with lightweight runtime probes to reconstruct an application’s I/O semantic intent during a single pre-execution pass. Guided by a large language model, Proteus makes layout decisions without requiring training or intrusive profiling. Evaluated on representative HPC workloads, Proteus achieves a layout decision accuracy of 91.30% and delivers speedups of up to 3.24× in write-intensive scenarios and 2.9× in metadata-intensive workloads.

Burst BufferData LayoutHPC Systems

DynamicAdaptiveClimb: Adaptive Cache Replacement with Dynamic Resizing

Nov 26, 2025
DB
Daniel Berend
🏛️ Ben Gurion University | Shiv Nadar Institution of Eminence

To address the declining cache hit ratio caused by dynamic shifts in cache access distributions, this paper proposes AdaptiveClimb, an adaptive cache replacement framework, and its extension DynamicAdaptiveClimb. Methodologically, it innovatively integrates oblivious statistical sampling with an incremental ranking advancement (CLIMB) mechanism to enable online, dynamic tuning of the jump distance; it further introduces, for the first time, a cache capacity self-scaling strategy that requires no complex state maintenance. The key contributions are a lightweight, low-overhead, fully unsupervised dual-adaptation mechanism—adapting both jump distance and cache size—which significantly enhances responsiveness to working-set fluctuations. Experimental evaluation across 1,067 real-world traces shows that DynamicAdaptiveClimb improves hit ratio by up to 29% over FIFO, outperforms state-of-the-art algorithms—including AdaptiveClimb and SIEVE—by 10–15%, and effectively reduces miss penalty.

Adapting cache promotion strategies based on access patternsAutomatically tuning cache size according to workload demandsOptimizing cache replacement policies for improved system performance

This work addresses the inefficiency of traditional database operators in remote memory environments, where ignoring communication round-trip overhead leads to poor performance during spilling. The paper presents the first operator optimization framework tailored for remote memory, incorporating the number of communication rounds into an operator-level latency cost model—thereby moving beyond the conventional I/O-centric optimization paradigm. It introduces round-trip-aware memory buffer partitioning strategies for key operators such as block nested-loop join, external merge sort, and external hash join. Implemented in DuckDB, the approach reduces communication rounds by up to 97% and operator execution time by up to 48% on a two-node platform. End-to-end evaluation on TPC-H and TPC-DS spilling queries demonstrates average speedups of 22.7% and 26.4%, respectively.

buffer allocationoperator optimizationout-of-memory processing

Hot Scholars

LB

Luca Benini

ETH Zürich, Università di Bologna
Integrated CircuitsComputer ArchitectureEmbedded SystemsVLSI
MV

Marian Verhelst

Micas - ESAT - KU Leuven, Belgium
Low-energy chip designsensor fusionmachine learningcross-layer optimization
YZ

Yuheng Zhao

Fudan University
Data VisualizationVisual AnalyticsHuman-AI Collaboration
JZ

Jieru Zhao

Associate Professor, Shanghai Jiao Tong University
Hardware-software co-designAI acceleration and systemCompilerFPGA
JZ

Jun Zhou

Ant Group, Alibaba Group, Zhejiang University
Distributed Machine LearningPrivacy Preserving Machine LearningGraph Neural NetworksAutoML