Score
Design and analyze abstract models of memory hierarchies (multiple cache and storage layers) that map access patterns and miss ratios to access cost and related metrics. Build analytical or simulation-based analyses to predict scalability, characterize operating regimes (e.g., power-law versus exponential miss behavior), and evaluate how hierarchy parameters affect latency, bandwidth, and resource use.
This work challenges the conventional assumption in time complexity analysis that ignores the impact of data movement costs in memory hierarchies as problem size scales. By leveraging an abstract memory hierarchy model and combining cache miss rate analysis with asymptotic complexity derivation, the study reveals for the first time that the data access cost for a broad class of common applications grows proportionally to \(N^{1/4}\) with input size \(N\). Furthermore, it distinguishes between scenarios where cache miss rates decay according to a power law versus an exponential law, and quantifies the resulting differences in scalability through constant factors. This refined characterization provides a more accurate performance prediction framework to guide algorithm design in modern memory-constrained architectures.
To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.
This work addresses the joint optimization of media selection, capacity allocation, and data placement (replication vs. tiering) for key-value caching across heterogeneous NVM/DRAM/disk storage under memory budget constraints. We introduce the first systematic modeling framework for multi-level non-volatile cache configurations, analytically characterize the operational regimes where replication or tiering dominates, and propose an adaptive configuration policy grounded in device failure rates and data update frequencies. Our methodology integrates cache access behavior modeling, hierarchical configuration optimization, and empirical validation using memcached benchmarks. Results demonstrate that tiering substantially outperforms replication under low device failure rates and high update workloads. Key contributions include: (1) a deployable, low-overhead configuration algorithm; (2) quantitative design guidelines for heterogeneous cache deployment; and (3) theoretical foundations for the reliability–performance trade-off in tiered caching systems.
The performance characteristics and architectural behaviors of cache-coherent interconnects—particularly Compute Express Link (CXL)—remain poorly understood in multi-vendor heterogeneous systems (e.g., CPU + CXL memory devices). Method: We construct a cross-vendor heterogeneous server cluster and propose Heimdall, the first fine-grained memory performance analysis framework tailored for CXL systems, accompanied by a lightweight microbenchmark suite. Through empirical measurement of CXL 3.0 protocol stack–hardware co-behavior, we systematically characterize memory latency, bandwidth, and coherence semantics across mainstream CXL devices. Contribution/Results: We uncover three previously unknown architectural blind spots and implicit protocol stack constraints. Leveraging these insights, we devise practical, workload-aware memory scheduling strategies for database and AI inference workloads. Our work provides both theoretical foundations and actionable guidelines for designing and optimizing cache-coherent heterogeneous systems.
To address the lack of a concurrency programming model ensuring data correctness and crash consistency in CXL-based disaggregated memory systems, this paper introduces CXL0—the first high-level programming model tailored for CXL. Our approach centers on three key contributions: (1) a formal operational semantics unifying memory sharing, persistence, and fault behaviors; (2) two general algorithmic transformation mechanisms—persistent linearization supporting partial failures, and a persistent algorithm restructuring framework resilient to full-system crashes; and (3) a prototype implementation of CXL0, including a hardware abstraction layer and preliminary performance evaluation. By bridging rigorous formal foundations with practical system design, CXL0 establishes the first theoretically sound and engineering-feasible foundation for building reliable concurrent programs on CXL platforms.
This work proposes a method to automatically derive efficient tiling and prefetching schedules from hardware cache hierarchies, eliminating reliance on empirical tuning. Building upon the Mathematics of Arrays framework, it introduces a machine shape—characterized by cache capacities, bandwidths, and occupancy sequences—to model multilevel caches, and designs novel operators that map this machine shape to tiling strategies, reducing unknown parameters to a small set of level-specific occupancies. The study addresses the fundamental open question of which parameters can be directly inferred from hardware specifications. Experiments on three real machines successfully reproduce expert-tuned tile sizes, and demonstrate that certain throughput-related parameters can indeed be derived from datasheets; however, cross-architecture portability remains limited and requires further improvement.
A lack of systematic methodologies for cross-platform performance and scalability comparison across heterogeneous HPC systems hinders fair and reproducible evaluation. Method: This paper proposes a unified cross-platform evaluation paradigm centered on the single compute node as the baseline unit. It integrates node-level performance measurement, weak/strong scaling analysis, and normalized metric comparison into a standardized experimental design, execution, and reporting workflow. A general-purpose validation framework is developed and empirically applied across diverse architectures—including CPUs, GPUs, and heterogeneous accelerators. Contributions/Results: (1) Establishes the single node as the minimal comparable unit for cross-platform assessment; (2) Provides a reusable, standardized evaluation template and integrated toolchain; (3) Significantly improves consistency, reproducibility, and interpretability of performance evaluation across heterogeneous HPC platforms.
This paper revisits and updates the 1987 “Five-Minute Rule” to reflect AI-era memory hierarchies dominated by GPUs and ultra-high-performance SSDs, addressing key limitations of the original rule—namely its neglect of host cost, physical constraints, and workload dynamics. Method: We propose a dynamic resource planning framework integrating DRAM bandwidth/capacity modeling, physics-aware SSD performance modeling, and workload-aware analysis. Leveraging the MQSim-Next simulator, we conduct sensitivity analysis to quantify cache threshold shifts under AI workloads. Contribution/Results: We demonstrate that the DRAM–NAND cache threshold has collapsed from minutes to seconds in AI scenarios—a first-time quantification. We establish NAND flash as a viable *active data layer*, enabling new hardware–software co-design paradigms. Two case studies validate the framework’s ability to expand the system design space, offering principled guidance for GPU-centric storage architecture.
This work addresses the inefficiency of large language models (LLMs), which rely on a monolithic context window as memory without hierarchical organization, leading to structural redundancy and resource waste. To overcome this limitation, the paper introduces the concept of virtual memory into LLM systems, proposing a novel L1–L3 multi-level memory architecture. A transparent proxy situated between the client and the inference API enables demand paging, page fault detection, and working-set page pinning. Integrated with dialogue compression and a page-fault-driven replacement policy, the system achieves up to a 93% reduction in context memory usage—from 5,038 KB to 339 KB—in real-world production settings, while offline simulations show a page fault rate of only 0.0254%, effectively transcending the constraints of conventional fixed-size context windows.