live memory hierarchy management

Designs and implements mechanisms and tools that coordinate live (hardware) and modeled/simulated memory hierarchies — including caches, TLBs, and main memory — to maintain consistency, control cross-host interference, and accelerate accurate memory-access modeling. Builds simulators, proxies, runtime managers, and analysis components that synchronize state between real and simulated memory structures and evaluate performance and correctness of memory-access behavior.

livememoryhierarchymanagement

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$206K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing cluster-level full-stack simulation struggles to simultaneously achieve high fidelity and high performance. This work proposes the concept of a “simulation-native operating system,” which integrates simulation control and orchestration into the OS kernel, thereby constructing a full-stack simulation framework built upon the Linux virtualization stack. The framework employs four key mechanisms—simulation-oriented scheduling, real-time memory hierarchy management, simulation-aware inter-process communication (IPC), and distributed simulation orchestration—to seamlessly co-execute real and simulated components without requiring modifications to production systems. Experimental results demonstrate that this approach significantly enhances the performance and configuration exploration efficiency of large-scale cluster simulations while preserving full-stack fidelity.

cluster-scale simulationdistributed systemsfull-stack fidelity

Traditional approaches struggle to provide deep visibility into the internal behavior of the gem5 simulator. This work proposes a non-intrusive, lightweight runtime call-stack analysis framework that, for the first time, treats the simulator’s own execution path as a novel lens for understanding simulated system behavior. Built upon the Linux perf_event interface, the framework enables parallel sampling, real-time symbol resolution, and hierarchical call-tree aggregation, with support for component-level customizable analysis. Experimental results demonstrate its effectiveness in uncovering performance bottlenecks in TimingSimpleCPU and identifying deadlock and livelock issues within the Ruby memory system—capturing critical behavioral characteristics that conventional statistical methods fail to detect.

cache coherencecall-stack profilinggem5

This work addresses the significant discrepancies between existing memory simulators and real hardware when predicting the performance of advanced memory systems, compounded by a lack of reliable validation methodologies. To tackle this issue, we propose the first multi-perspective co-validation framework that systematically evaluates simulation accuracy from three complementary dimensions: the memory simulator itself, the CPU–memory interface, and application-level behavior. Our analysis reveals that inaccuracies at the interface layer are a primary source of simulation distortion. Building on this insight, we integrate mainstream simulators—Ramulator, Ramulator2, and DRAMsim3—into the ZSim platform and implement targeted corrections and enhancements at the interface layer. Experimental results demonstrate that the refined simulators achieve substantially improved fidelity across diverse workloads, yielding predictions that closely align with real-system performance.

CPU-memory interfacehardware validationmemory simulation

Modal Abstractions for Virtualizing Memory Addresses

Jul 26, 2023
IK
Ismail Kuru
🏛️ Drexel University

Formal verification of operating system kernel virtual memory management (VMM) code remains challenging due to hardware interface complexity and difficulties in semantically modeling dynamic multi-address-space switching. This paper addresses these challenges by introducing a modal-logic-based abstraction of address spaces. Our method features: (1) a novel modal assertion ([r]P) to express truth relative to an address space (r); (2) a precise virtual *points-to* relation that faithfully models hardware page-table translation semantics; and (3) the first fully mechanized formal verification—within the Iris separation logic framework and Coq—supporting instruction sequences spanning multiple address spaces. We verify critical VMM operations including address-space switching and page-table updates. All semantic definitions and proofs are entirely mechanized in Coq, achieving significantly stronger verification guarantees than prior approaches.

Enabling modal assertions for multiple address space verificationHandling hardware interface challenges in OS kernelsVerifying virtual memory management code complexity

A Programming Model for Disaggregated Memory over CXL

Jul 23, 2024
GA
Gal Assa
🏛️ Technion | ETH Zurich | Tel Aviv University

To address the lack of a concurrency programming model ensuring data correctness and crash consistency in CXL-based disaggregated memory systems, this paper introduces CXL0—the first high-level programming model tailored for CXL. Our approach centers on three key contributions: (1) a formal operational semantics unifying memory sharing, persistence, and fault behaviors; (2) two general algorithmic transformation mechanisms—persistent linearization supporting partial failures, and a persistent algorithm restructuring framework resilient to full-system crashes; and (3) a prototype implementation of CXL0, including a hardware abstraction layer and preliminary performance evaluation. By bridging rigorous formal foundations with practical system design, CXL0 establishes the first theoretically sound and engineering-feasible foundation for building reliable concurrent programs on CXL platforms.

CXL environmentdata correctness and reliabilityprogramming methods

Latest Papers

What's happening recently
View more

This work addresses the limited design space of existing Processing-in-Memory (PIM) simulators, which struggle to support diverse memory technologies, flexible processing element (PE) deployment, and end-to-end evaluation. To overcome these limitations, we present PIMID—the first full-system PIM simulator within a unified framework—integrating execution-driven and trace-driven methodologies. PIMID supports eleven memory technologies (including DRAM, SRAM, and non-volatile memories), configurable PE placement and scale, and compatibility with both OpenMP shared-memory and MPI message-passing programming models. It provides fine-grained latency and energy breakdowns and features a YAML-based plugin mechanism for future extensibility. Experimental results reveal that memory technology impacts performance by over an order of magnitude, with the best conventional main memory not necessarily optimal as a PIM substrate; regular kernels exhibit superlinear performance scaling with PE count; graph traversal is bottlenecked by MPI communication; and shared-memory offloading on HBM3 achieves both energy efficiency and end-to-end speedup.

design space explorationexecution modelfull-system simulation

Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.

distributed systemsparallel systemsruntime control

Existing distributed LLM serving simulators lack closed-loop execution and hardware-free profiling for timing prediction. This work proposes a closed-loop simulation framework that integrates specification-driven analytical timing modeling with a stateful serving loop. It leverages iSTAGE to generate profile-free traces and decouples component ownership to support cross-platform portability. Compatible with the vLLM interface, the framework enables unmodified benchmarks to run directly while accurately capturing scheduling, queuing, and KV cache feedback. Experiments demonstrate steady-state throughput and multi-turn performance errors of only 3.6% and 9.9%, respectively. Furthermore, the study reveals novel mechanisms, including an inversion in the HBM bandwidth-capacity trade-off and a concurrency-induced bottleneck shift from memory constraints to scheduling overhead.

Architecture explorationBenchmark validationDistributed LLM serving

This work addresses the challenges of efficiency, scalability, and consistency in distributed management of KV caches for large language model (LLM) services. It proposes the first four-dimensional taxonomy—spanning locality, lifetime, ownership, and storage substrate—to systematically analyze over 30 existing studies, identifying five architectural paradigms: local paging, decoupled pipelining, shared storage, memory pooling, and hybrid hierarchical designs. The study reveals that “ownership” is a key differentiator in distributed KV cache architectures and highlights the absence of seven KV-specific metrics in current evaluation methodologies. Furthermore, it connects these gaps to six critical open problems, including fault tolerance, isolation, hierarchical eviction, and speculative decoding.

cache managementdistributed memoryKV cache

This work addresses the inefficiency of frequent memory accesses in large language model (LLM) inference, which underutilizes the substantial last-level cache capacity of modern multi-core CPUs. To overcome this limitation, the authors propose a cache-resident execution model that decouples weight-intensive computations from attention mechanisms and KV cache management, assigning them to dedicated resource domains. By relaxing synchronization constraints based on sub-operator dependencies, the approach breaks conventional operator boundaries through weight cache residency decoupled from KV cache capacity, locality-aware data placement, and a lightweight static runtime. This design significantly reduces coordination overhead, achieving 2.04× to 11.51× per-token inference speedup on Llama-3.2-3B and Llama-2-7B models, with a theoretical peak acceleration of 13.9×.

cache residencyKV-cacheLLM inference

Hot Scholars

PZ

Pengfei Zuo

Huawei
AI InfrastructureCloud InfrastructureMachine Learning SystemsMemory Systems
TR

Tajana Rosing

Distinguished Professor, UCSD
computer architecturecyber-physical systemssystem energy efficiency
SD

Sheng Di

Argonne National Labratory, IEEE Senior Member
HPCData CompressionResilienceCloud/Grid Computing/P2P
YL

Yongpan Liu

Professor @ Tsinghua University
Machine LearningNonvolatile Memory and ComputingEnergy Efficient VLSIEmbedded System
MD

Mario Di Francesco

Department of Computer Science, Aalto University
Wireless networkingmobile computingInternet of Things