Score
Designs and implements buffer allocation and management systems that provision and resize buffer capacity, orchestrate movement of data between storage and compute, and implement cache eviction and replacement policies. Analyzes local data reuse patterns and access behavior to optimize placement, lifetime, and transfer scheduling of buffered data to reduce latency and memory footprint.
The widening gap between CPU and memory access latency poses a fundamental challenge to database performance. Method: This paper systematically surveys four decades of buffer management evolution, analyzing classical algorithms (e.g., LRU-K, 2Q, LIRS, ARC), ML-enhanced approaches, and memory-disaggregated architectures (NVM-aware hierarchies, RDMA-decoupled designs). It introduces the first holistic, time-spanning analytical framework for buffer management evolution and proposes a novel cross-layer adaptive paradigm integrating ML-driven policy learning with eBPF-based kernel extensibility. Contribution/Results: Grounded in empirical analysis of 50+ top-tier conference papers and industrial systems (Linux, PostgreSQL, Oracle), the work distills core design trade-offs and identifies key challenges—including cache coherence across heterogeneous memory tiers, low-overhead decision-making, and OS–DBMS co-design—along with concrete research directions toward scalable, adaptive buffer management.
To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.
To address the challenge of optimizing data placement in HBM+DDR heterogeneous memory systems, this paper proposes a lightweight, non-intrusive, application-level memory analysis and control framework. Leveraging a detailed memory subsystem model on Intel Sapphire Rapids platforms—integrated with empirical bandwidth/latency measurements, runtime allocation tracing, and policy injection—the work quantifies, for the first time, the performance inflection point of HBM-DRAM co-utilization: retaining only 60–75% of critical data in HBM achieves 90% of the platform’s peak performance. The approach requires no source-code modification or recompilation. Evaluated across multiple benchmarks, it significantly improves performance for memory-intensive applications while reducing HBM resource consumption by over 30%. This establishes a deployable, fine-grained data placement optimization paradigm for heterogeneous memory systems.
This work addresses the challenges of efficiency, scalability, and consistency in distributed management of KV caches for large language model (LLM) services. It proposes the first four-dimensional taxonomy—spanning locality, lifetime, ownership, and storage substrate—to systematically analyze over 30 existing studies, identifying five architectural paradigms: local paging, decoupled pipelining, shared storage, memory pooling, and hybrid hierarchical designs. The study reveals that “ownership” is a key differentiator in distributed KV cache architectures and highlights the absence of seven KV-specific metrics in current evaluation methodologies. Furthermore, it connects these gaps to six critical open problems, including fault tolerance, isolation, hierarchical eviction, and speculative decoding.
This work addresses the joint optimization of media selection, capacity allocation, and data placement (replication vs. tiering) for key-value caching across heterogeneous NVM/DRAM/disk storage under memory budget constraints. We introduce the first systematic modeling framework for multi-level non-volatile cache configurations, analytically characterize the operational regimes where replication or tiering dominates, and propose an adaptive configuration policy grounded in device failure rates and data update frequencies. Our methodology integrates cache access behavior modeling, hierarchical configuration optimization, and empirical validation using memcached benchmarks. Results demonstrate that tiering substantially outperforms replication under low device failure rates and high update workloads. Key contributions include: (1) a deployable, low-overhead configuration algorithm; (2) quantitative design guidelines for heterogeneous cache deployment; and (3) theoretical foundations for the reliability–performance trade-off in tiered caching systems.
Traditional virtual memory–assisted buffer management struggles to efficiently support data migration and access in multi-tier memory architectures. This work proposes vmcacheⁿ, a framework that extends the conventional two-level (DRAM–disk) buffering mechanism to an n-level hierarchy (DRAM–remote memory–disk). Leveraging the operating system’s virtual memory subsystem and page migration facilities, vmcacheⁿ constructs a multi-tier cache pool by integrating remote memory technologies such as NUMA and CXL. To enable fine-grained and low-overhead cross-tier page migration, the framework introduces a new system call, move_pages2. Experimental evaluation under the TPC-C workload demonstrates that vmcacheⁿ achieves up to a 4× improvement in query throughput compared to the original vmcache.
Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.
Existing burst buffer file systems suffer from performance degradation due to rigid, fixed data layouts that fail to adapt to diverse application I/O behaviors. This work proposes Proteus, a system that treats data layout as a first-class optimization target by integrating static code analysis with lightweight runtime probes to reconstruct an application’s I/O semantic intent during a single pre-execution pass. Guided by a large language model, Proteus makes layout decisions without requiring training or intrusive profiling. Evaluated on representative HPC workloads, Proteus achieves a layout decision accuracy of 91.30% and delivers speedups of up to 3.24× in write-intensive scenarios and 2.9× in metadata-intensive workloads.
To address the declining cache hit ratio caused by dynamic shifts in cache access distributions, this paper proposes AdaptiveClimb, an adaptive cache replacement framework, and its extension DynamicAdaptiveClimb. Methodologically, it innovatively integrates oblivious statistical sampling with an incremental ranking advancement (CLIMB) mechanism to enable online, dynamic tuning of the jump distance; it further introduces, for the first time, a cache capacity self-scaling strategy that requires no complex state maintenance. The key contributions are a lightweight, low-overhead, fully unsupervised dual-adaptation mechanism—adapting both jump distance and cache size—which significantly enhances responsiveness to working-set fluctuations. Experimental evaluation across 1,067 real-world traces shows that DynamicAdaptiveClimb improves hit ratio by up to 29% over FIFO, outperforms state-of-the-art algorithms—including AdaptiveClimb and SIEVE—by 10–15%, and effectively reduces miss penalty.
This work addresses the inefficiency of traditional database operators in remote memory environments, where ignoring communication round-trip overhead leads to poor performance during spilling. The paper presents the first operator optimization framework tailored for remote memory, incorporating the number of communication rounds into an operator-level latency cost model—thereby moving beyond the conventional I/O-centric optimization paradigm. It introduces round-trip-aware memory buffer partitioning strategies for key operators such as block nested-loop join, external merge sort, and external hash join. Implemented in DuckDB, the approach reduces communication rounds by up to 97% and operator execution time by up to 48% on a two-node platform. End-to-end evaluation on TPC-H and TPC-DS spilling queries demonstrates average speedups of 22.7% and 26.4%, respectively.