Score
Designs, implements, and evaluates hardware and software components that store, organize, access, and manage data in computing systems, including volatile and nonvolatile storage, caches, memory allocation, garbage collection, and memory hierarchies. Works on performance, correctness, consistency, persistence, reliability, and interfaces between processors and memory.
This work addresses the joint optimization of media selection, capacity allocation, and data placement (replication vs. tiering) for key-value caching across heterogeneous NVM/DRAM/disk storage under memory budget constraints. We introduce the first systematic modeling framework for multi-level non-volatile cache configurations, analytically characterize the operational regimes where replication or tiering dominates, and propose an adaptive configuration policy grounded in device failure rates and data update frequencies. Our methodology integrates cache access behavior modeling, hierarchical configuration optimization, and empirical validation using memcached benchmarks. Results demonstrate that tiering substantially outperforms replication under low device failure rates and high update workloads. Key contributions include: (1) a deployable, low-overhead configuration algorithm; (2) quantitative design guidelines for heterogeneous cache deployment; and (3) theoretical foundations for the reliability–performance trade-off in tiered caching systems.
To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.
This paper addresses the challenge of verifying concurrent programs arising from ambiguous formal specifications of weak memory models. It systematically surveys and comparatively analyzes two dominant modeling paradigms—operational semantics and axiomatic semantics—and introduces, for the first time, a unified framework for constructing execution traces and analyzing memory event relations. Using a simplified x86 model as a case study, the work integrates hardware microarchitectural features, computability theory, and advances in verification tools to achieve precise alignment between formal models and real hardware behavior. Its contributions are threefold: (1) establishing a cross-paradigm formal benchmark for semantic comparison; (2) proposing a unified analysis framework that balances verifiability and interpretability; and (3) delivering a comprehensive landscape of formal foundations and tooling support for safety-critical low-level software, thereby fostering synergistic advancement of theoretical rigor and engineering practicality.
To address critical challenges in SoC design—including ambiguous system-level modeling semantics, poor interoperability across heterogeneous computational models (e.g., dataflow and neural networks), and the decoupling of design-space exploration from verification—this paper proposes a co-communication mechanism ensuring semantic consistency across multiple models. The approach establishes an integrated toolchain supporting system-level modeling, simulation-driven verification, hardware-software co-design space exploration, and joint power-performance analysis. Innovatively, it unifies dataflow modeling with system-level abstractions to enable functional correctness verification and quantitative energy-efficiency evaluation for representative applications such as video processing and AI acceleration. Experimental results demonstrate that the methodology significantly improves early-stage SoC design iteration efficiency and enhances the reliability of architectural decision-making.
To address performance, energy-efficiency, and latency bottlenecks arising from data movement between main memory and processors in data-intensive computing, this paper proposes and systematically advances the practical deployment of Processing-in-Memory (PIM). We introduce a unified PIM architecture framework that innovatively integrates 3D-stacked logic layers, on-die in-memory computing units, memory controller–integrated accelerators, on-chip ECC-enhanced DRAM, and hardware-level RowHammer mitigation. This holistic design significantly reduces data movement overhead, achieving energy-efficiency improvements of several-fold to over an order of magnitude across representative workloads, while ensuring scalability and reliability. The framework establishes a foundation for low-power, high-throughput memory-compute infrastructure applicable to both server and mobile platforms. By bridging architectural innovation with system-level implementation, our work accelerates the transition of PIM from theoretical concept to real-world deployment.
This work addresses the significant discrepancies between existing memory simulators and real hardware when predicting the performance of advanced memory systems, compounded by a lack of reliable validation methodologies. To tackle this issue, we propose the first multi-perspective co-validation framework that systematically evaluates simulation accuracy from three complementary dimensions: the memory simulator itself, the CPU–memory interface, and application-level behavior. Our analysis reveals that inaccuracies at the interface layer are a primary source of simulation distortion. Building on this insight, we integrate mainstream simulators—Ramulator, Ramulator2, and DRAMsim3—into the ZSim platform and implement targeted corrections and enhancements at the interface layer. Experimental results demonstrate that the refined simulators achieve substantially improved fidelity across diverse workloads, yielding predictions that closely align with real-system performance.
Existing methodologies—such as SPEC CPU2017, Design of Experiments (DoE), and Randomized Controlled Trials (RCTs)—struggle to accurately attribute overall system performance to individual hardware components due to their inability to effectively isolate component-level contributions, resulting in substantial evaluation variability (SPEC score deviations ranging from 12.16% to 436.80%). This work proposes a novel methodology that integrates controlled experimentation with a theoretical attribution model, enabling, for the first time, precise and stable attribution of system performance to specific hardware components. The proposed approach significantly outperforms conventional techniques, offering high cost-effectiveness while overcoming inherent limitations in component evaluation and system design present in current practices.
Modern computing systems face fundamental bottlenecks in memory safety and efficiency due to the absence of native hardware/software interface support for high-level semantic properties—such as object identity, boundaries, and lifetime. To address this, we propose the *object-aware memory* paradigm, elevating descriptors to first-class hardware abstractions and establishing a novel descriptor-based addressing taxonomy and unified analytical framework. Our approach integrates an object-aware memory model, hybrid tagged-pointer encoding, and cross-layer semantic propagation to enable dynamic object identification, boundary enforcement, and coordinated lifetime management. We implement and evaluate CentroID, a prototype system demonstrating substantial improvements in security (e.g., eliminating spatial/temporal memory errors), performance (near-native execution overhead), and practical deployability. This work establishes the first structured research framework for object-aware memory and provides both theoretical foundations and concrete architectural pathways toward semantics-driven next-generation memory systems.
Accurately modeling and quantitatively evaluating performance bottlenecks in real-world systems remains challenging. Method: This paper proposes a theory-driven, practice-oriented, progressive performance modeling framework that integrates queuing theory, Markov models, load-testing-based modeling, and system simulation. It employs a three-tiered problem design—foundational modeling → dynamic workload analysis → industrial-scale system simulation—to enable capability transfer from classroom training to complex system analysis. Contribution/Results: The framework innovatively couples quantitative modeling techniques with hierarchical pedagogical practices, establishing a scalable, verifiable, integrated teaching–practice ecosystem for performance evaluation. Experimental results demonstrate significant improvements in learners’ modeling accuracy and solution efficiency for large-scale system performance problems; the framework has been successfully deployed in multiple industrial system performance optimization scenarios.
This work addresses the gap in current systems education, where learning resources often consist of superficial tutorials or AI-generated summaries that inadequately convey foundational design principles and thus fail to cultivate robust engineering capabilities. To remedy this, we propose a structured learning pathway centered on seminal research papers from distributed systems, operating systems, and big data domains. Integrating insights from leading academic curricula and industry practices, our approach emphasizes technical depth and problem-solving reasoning. By engaging learners in close reading of original literature, critical analysis of architectural trade-offs, and cross-domain synthesis, the framework fosters a deep understanding of underlying mechanisms and cultivates systems thinking—thereby equipping practitioners to effectively tackle complex engineering challenges and progress toward professional-level systems expertise.