Score
Designs and implements mechanisms and tools that coordinate live (hardware) and modeled/simulated memory hierarchies — including caches, TLBs, and main memory — to maintain consistency, control cross-host interference, and accelerate accurate memory-access modeling. Builds simulators, proxies, runtime managers, and analysis components that synchronize state between real and simulated memory structures and evaluate performance and correctness of memory-access behavior.
Existing cluster-level full-stack simulation struggles to simultaneously achieve high fidelity and high performance. This work proposes the concept of a “simulation-native operating system,” which integrates simulation control and orchestration into the OS kernel, thereby constructing a full-stack simulation framework built upon the Linux virtualization stack. The framework employs four key mechanisms—simulation-oriented scheduling, real-time memory hierarchy management, simulation-aware inter-process communication (IPC), and distributed simulation orchestration—to seamlessly co-execute real and simulated components without requiring modifications to production systems. Experimental results demonstrate that this approach significantly enhances the performance and configuration exploration efficiency of large-scale cluster simulations while preserving full-stack fidelity.
Traditional approaches struggle to provide deep visibility into the internal behavior of the gem5 simulator. This work proposes a non-intrusive, lightweight runtime call-stack analysis framework that, for the first time, treats the simulator’s own execution path as a novel lens for understanding simulated system behavior. Built upon the Linux perf_event interface, the framework enables parallel sampling, real-time symbol resolution, and hierarchical call-tree aggregation, with support for component-level customizable analysis. Experimental results demonstrate its effectiveness in uncovering performance bottlenecks in TimingSimpleCPU and identifying deadlock and livelock issues within the Ruby memory system—capturing critical behavioral characteristics that conventional statistical methods fail to detect.
This work addresses the significant discrepancies between existing memory simulators and real hardware when predicting the performance of advanced memory systems, compounded by a lack of reliable validation methodologies. To tackle this issue, we propose the first multi-perspective co-validation framework that systematically evaluates simulation accuracy from three complementary dimensions: the memory simulator itself, the CPU–memory interface, and application-level behavior. Our analysis reveals that inaccuracies at the interface layer are a primary source of simulation distortion. Building on this insight, we integrate mainstream simulators—Ramulator, Ramulator2, and DRAMsim3—into the ZSim platform and implement targeted corrections and enhancements at the interface layer. Experimental results demonstrate that the refined simulators achieve substantially improved fidelity across diverse workloads, yielding predictions that closely align with real-system performance.
Formal verification of operating system kernel virtual memory management (VMM) code remains challenging due to hardware interface complexity and difficulties in semantically modeling dynamic multi-address-space switching. This paper addresses these challenges by introducing a modal-logic-based abstraction of address spaces. Our method features: (1) a novel modal assertion ([r]P) to express truth relative to an address space (r); (2) a precise virtual *points-to* relation that faithfully models hardware page-table translation semantics; and (3) the first fully mechanized formal verification—within the Iris separation logic framework and Coq—supporting instruction sequences spanning multiple address spaces. We verify critical VMM operations including address-space switching and page-table updates. All semantic definitions and proofs are entirely mechanized in Coq, achieving significantly stronger verification guarantees than prior approaches.
To address the lack of a concurrency programming model ensuring data correctness and crash consistency in CXL-based disaggregated memory systems, this paper introduces CXL0—the first high-level programming model tailored for CXL. Our approach centers on three key contributions: (1) a formal operational semantics unifying memory sharing, persistence, and fault behaviors; (2) two general algorithmic transformation mechanisms—persistent linearization supporting partial failures, and a persistent algorithm restructuring framework resilient to full-system crashes; and (3) a prototype implementation of CXL0, including a hardware abstraction layer and preliminary performance evaluation. By bridging rigorous formal foundations with practical system design, CXL0 establishes the first theoretically sound and engineering-feasible foundation for building reliable concurrent programs on CXL platforms.
This work addresses the limited design space of existing Processing-in-Memory (PIM) simulators, which struggle to support diverse memory technologies, flexible processing element (PE) deployment, and end-to-end evaluation. To overcome these limitations, we present PIMID—the first full-system PIM simulator within a unified framework—integrating execution-driven and trace-driven methodologies. PIMID supports eleven memory technologies (including DRAM, SRAM, and non-volatile memories), configurable PE placement and scale, and compatibility with both OpenMP shared-memory and MPI message-passing programming models. It provides fine-grained latency and energy breakdowns and features a YAML-based plugin mechanism for future extensibility. Experimental results reveal that memory technology impacts performance by over an order of magnitude, with the best conventional main memory not necessarily optimal as a PIM substrate; regular kernels exhibit superlinear performance scaling with PE count; graph traversal is bottlenecked by MPI communication; and shared-memory offloading on HBM3 achieves both energy efficiency and end-to-end speedup.
Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.
Existing distributed LLM serving simulators lack closed-loop execution and hardware-free profiling for timing prediction. This work proposes a closed-loop simulation framework that integrates specification-driven analytical timing modeling with a stateful serving loop. It leverages iSTAGE to generate profile-free traces and decouples component ownership to support cross-platform portability. Compatible with the vLLM interface, the framework enables unmodified benchmarks to run directly while accurately capturing scheduling, queuing, and KV cache feedback. Experiments demonstrate steady-state throughput and multi-turn performance errors of only 3.6% and 9.9%, respectively. Furthermore, the study reveals novel mechanisms, including an inversion in the HBM bandwidth-capacity trade-off and a concurrency-induced bottleneck shift from memory constraints to scheduling overhead.
This work addresses the challenges of efficiency, scalability, and consistency in distributed management of KV caches for large language model (LLM) services. It proposes the first four-dimensional taxonomy—spanning locality, lifetime, ownership, and storage substrate—to systematically analyze over 30 existing studies, identifying five architectural paradigms: local paging, decoupled pipelining, shared storage, memory pooling, and hybrid hierarchical designs. The study reveals that “ownership” is a key differentiator in distributed KV cache architectures and highlights the absence of seven KV-specific metrics in current evaluation methodologies. Furthermore, it connects these gaps to six critical open problems, including fault tolerance, isolation, hierarchical eviction, and speculative decoding.
This work addresses the inefficiency of frequent memory accesses in large language model (LLM) inference, which underutilizes the substantial last-level cache capacity of modern multi-core CPUs. To overcome this limitation, the authors propose a cache-resident execution model that decouples weight-intensive computations from attention mechanisms and KV cache management, assigning them to dedicated resource domains. By relaxing synchronization constraints based on sub-operator dependencies, the approach breaks conventional operator boundaries through weight cache residency decoupled from KV cache capacity, locality-aware data placement, and a lightweight static runtime. This design significantly reduces coordination overhead, achieving 2.04× to 11.51× per-token inference speedup on Llama-3.2-3B and Llama-2-7B models, with a theoretical peak acceleration of 13.9×.