Score
Designs, builds, and analyzes memory systems, architectures, and mechanisms to reduce memory footprint and improve access efficiency, bandwidth, and latency across memory hierarchies. Develops and evaluates memory layouts, allocation and management strategies (including adaptive and self‑optimizing policies), bandwidth tuning techniques, and measurement tools for memory footprint and bandwidth analysis.
To address the challenge of optimizing data placement in HBM+DDR heterogeneous memory systems, this paper proposes a lightweight, non-intrusive, application-level memory analysis and control framework. Leveraging a detailed memory subsystem model on Intel Sapphire Rapids platforms—integrated with empirical bandwidth/latency measurements, runtime allocation tracing, and policy injection—the work quantifies, for the first time, the performance inflection point of HBM-DRAM co-utilization: retaining only 60–75% of critical data in HBM achieves 90% of the platform’s peak performance. The approach requires no source-code modification or recompilation. Evaluated across multiple benchmarks, it significantly improves performance for memory-intensive applications while reducing HBM resource consumption by over 30%. This establishes a deployable, fine-grained data placement optimization paradigm for heterogeneous memory systems.
With SRAM/DRAM costs plateauing, memory has become the primary bottleneck for system cost and energy efficiency. Method: This paper proposes a memory specialization paradigm that departs from conventional hierarchical memory architectures. It introduces two application-tailored memory classes: Long-term RAM (LtRAM) for long-lived, read-intensive data, and Short-term RAM (StRAM) for short-lived, high-frequency transient accesses. This approach requires explicit OS-level management to enable non-hierarchical memory resource scheduling. Leveraging emerging memory technologies, the paper designs hardware architectures, system interfaces, and integration mechanisms for LtRAM and StRAM. Contribution/Results: The work identifies key technical challenges—including coherence, migration, and interface standardization—and outlines concrete implementation pathways. By rethinking memory as a heterogeneous, application-aware resource rather than a monolithic hierarchy, it establishes a foundational architectural direction for efficient, scalable computing systems.
Memory latency, bandwidth, capacity, and energy consumption have become critical bottlenecks for large-scale parallel systems. This paper proposes a decentralized hierarchical memory architecture: ultra-large (TB–PB scale) memory is partitioned into small, compute-tightly-coupled nodes; leveraging 2.5D/3D integration, high-speed caches and DRAM main memory are co-packaged to realize a near-compute memory paradigm. Crucially, the architecture introduces hardware-explicit mechanisms for managing both memory capacity and physical distance, enabling fine-grained software control over data placement and migration. This design significantly reduces memory access latency and dynamic power consumption, improves bandwidth utilization, and mitigates signal integrity and scalability challenges inherent in conventional memory systems. The result is a novel, energy-efficient, and scalable memory architecture tailored for next-generation large-scale computing systems.
Modern computing systems face an increasingly severe memory bottleneck—characterized by high energy consumption, performance limitations, reduced reliability, escalating costs, and substantial hardware overhead—particularly constraining data-intensive workloads such as AI and graph processing. To address this, we propose a memory-centric design and execution paradigm that transcends the traditional processor-centric architecture’s passive reliance on memory. Our approach introduces a novel memory-autonomous management mechanism synergistically integrated with Processing-in-Memory (PIM), combining adaptive in-memory control circuits, near-memory and in-memory computing architectures, cross-layer hardware-software co-scheduling, and evolutionary migration strategies. This transforms memory from a “dumb storage” unit into an intelligent, programmable computing resource. Experimental results demonstrate order-of-magnitude improvements in system energy efficiency and performance, along with significant mitigation of reliability threats such as RowHammer. The framework establishes a scalable, low-power foundation for next-generation data-intensive applications.
This work addresses the joint optimization of media selection, capacity allocation, and data placement (replication vs. tiering) for key-value caching across heterogeneous NVM/DRAM/disk storage under memory budget constraints. We introduce the first systematic modeling framework for multi-level non-volatile cache configurations, analytically characterize the operational regimes where replication or tiering dominates, and propose an adaptive configuration policy grounded in device failure rates and data update frequencies. Our methodology integrates cache access behavior modeling, hierarchical configuration optimization, and empirical validation using memcached benchmarks. Results demonstrate that tiering substantially outperforms replication under low device failure rates and high update workloads. Key contributions include: (1) a deployable, low-overhead configuration algorithm; (2) quantitative design guidelines for heterogeneous cache deployment; and (3) theoretical foundations for the reliability–performance trade-off in tiered caching systems.
This work addresses the challenge of achieving fine-grained, portable, and rapidly responsive memory bandwidth regulation in real-time multicore systems. It proposes the first hardware-assisted memory bandwidth regulator leveraging Arm’s CoreSight Embedded Trace Macrocell (ETM) to enable interrupt-driven, per-core bandwidth control with microsecond-level temporal resolution. The approach incurs minimal software overhead and effectively bridges the gap between MemPol—known for high precision—and MemGuard—valued for its portability—while supporting novel regulation policies. Experimental evaluations across multiple 64-bit Arm platforms, including the Zynq UltraScale+, demonstrate that the proposed solution outperforms existing methods in terms of effectiveness, scalability, and regulation accuracy.
This work addresses the growing challenge posed by the expanding data footprint of modern applications, which renders memory systems a critical bottleneck for both performance and energy efficiency—constraints that traditional microarchitectures struggle to overcome. To this end, the paper introduces a data-driven microarchitectural design paradigm that systematically integrates lightweight machine learning with application-specific data semantic features across multiple processor components. Key contributions include a reinforcement learning–based hardware prefetcher, a perceptron-driven off-chip access predictor, a synergistic mechanism coordinating prefetching and prediction, and a predictability-aware memory access elimination technique leveraging both address and value repetition. Experimental results demonstrate that the proposed approach substantially outperforms state-of-the-art solutions, delivering significant improvements in both performance and energy efficiency.
This work addresses the significant yet underexplored impact of parameter configuration on performance in memory tiering systems, where efficient automated tuning mechanisms are lacking. The authors propose PTMT, a lightweight framework that systematically categorizes memory tiering parameters and reveals their performance sensitivity for the first time. PTMT introduces a hybrid adaptive tuning mechanism that combines offline performance profiling with online reinforcement learning to achieve high-efficacy optimization at low overhead. By co-designing memory access profiling and page migration, PTMT improves performance by 30%, 26%, 21%, and 14% over TPP, UPM, Colloid, and AutoNUMA, respectively, and outperforms the best existing approaches by 32% on average.
This work addresses the significant discrepancies between existing memory simulators and real hardware when predicting the performance of advanced memory systems, compounded by a lack of reliable validation methodologies. To tackle this issue, we propose the first multi-perspective co-validation framework that systematically evaluates simulation accuracy from three complementary dimensions: the memory simulator itself, the CPU–memory interface, and application-level behavior. Our analysis reveals that inaccuracies at the interface layer are a primary source of simulation distortion. Building on this insight, we integrate mainstream simulators—Ramulator, Ramulator2, and DRAMsim3—into the ZSim platform and implement targeted corrections and enhancements at the interface layer. Experimental results demonstrate that the refined simulators achieve substantially improved fidelity across diverse workloads, yielding predictions that closely align with real-system performance.
Traditional memory systems rely on static, handcrafted heuristics that struggle to adapt to dynamic workloads. This work presents the first systematic integration of lightweight machine learning techniques into the memory subsystem, introducing an adaptive, data-driven self-optimizing architecture. The proposed framework comprises Pythia, a reinforcement learning–based prefetcher; Hermes, a perceptron-driven off-chip bandwidth predictor; and Sibyl, a reinforcement learning–guided data placement policy. All three components consistently outperform state-of-the-art hand-tuned designs, delivering substantial improvements in both system performance and energy efficiency while incurring only modest hardware overhead.