Score
Designs, implements, and analyzes network infrastructures and persistent data storage systems, covering network architectures, protocols, and equipment as well as storage architectures (block/file/object), distributed and local storage, and storage networking. Focuses on performance, scalability, availability, security, and the interactions between networking and storage such as data transfer, replication, caching, and storage protocols.
To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.
This work addresses the lack of systematic educational resources in high-performance computing (HPC) networking, which poses a significant barrier for researchers entering the field. It presents the first comprehensive integration of the HPC networking stack, covering communication protocol layers, programming interfaces such as MPI, control plane mechanisms, high-speed interconnect technologies, and custom link-layer hardware. The exposition is anchored by a detailed case study of the El Capitan supercomputer architecture at Lawrence Livermore National Laboratory. By offering a well-structured, practice-oriented primer, this contribution fills a critical educational gap and substantially lowers the entry barrier for researchers seeking to master core HPC networking technologies.
Modern data storage systems suffer from latent cross-layer faults due to tight hardware–software coupling across multiple abstraction layers, often leading to silent data corruption or unrecoverable data loss. To address this, we propose the first cross-layer fault-tolerance analysis framework targeting heterogeneous storage stacks—including SSDs, persistent memory, local file systems, and distributed storage. Our approach combines architectural modeling of the full stack, systematic injection of representative defects, and precise tracking of fault propagation across hardware–firmware–software boundaries to expose error propagation paths and consistency violation mechanisms. Through empirical evaluation across widely deployed systems, we identify critical vulnerabilities impacting data integrity and quantify coverage gaps in existing fault-tolerance techniques. The framework provides a scalable, principled methodology for analyzing cross-layer resilience and establishes concrete, actionable directions for designing next-generation highly reliable storage systems.
This work addresses the joint optimization of media selection, capacity allocation, and data placement (replication vs. tiering) for key-value caching across heterogeneous NVM/DRAM/disk storage under memory budget constraints. We introduce the first systematic modeling framework for multi-level non-volatile cache configurations, analytically characterize the operational regimes where replication or tiering dominates, and propose an adaptive configuration policy grounded in device failure rates and data update frequencies. Our methodology integrates cache access behavior modeling, hierarchical configuration optimization, and empirical validation using memcached benchmarks. Results demonstrate that tiering substantially outperforms replication under low device failure rates and high update workloads. Key contributions include: (1) a deployable, low-overhead configuration algorithm; (2) quantitative design guidelines for heterogeneous cache deployment; and (3) theoretical foundations for the reliability–performance trade-off in tiered caching systems.
Conventional network telemetry frameworks struggle to support fine-grained traffic measurement, performance diagnostics, and attack detection under stringent memory and computational constraints of high-speed network devices. Method: This paper proposes a lightweight, real-time online telemetry framework that systematically integrates compact data structures—including Bloom filter variants, Count-Min Sketch, and HyperLogLog—with streaming algorithms, hierarchical sampling, and P4-programmable data-plane co-design to comply with hardware limitations. Contribution/Results: Evaluated at line rate exceeding 100 Gbps, the framework reduces memory footprint by over 60% compared to state-of-the-art approaches while maintaining sub-1% flow frequency estimation error. It achieves an optimal trade-off among accuracy, throughput, and resource overhead, thereby significantly enhancing the feasibility and practicality of telemetry in high-bandwidth environments.
This study addresses the critical impact of content caching efficiency on content delivery performance in Named Data Networking (NDN) by providing a systematic survey of caching algorithms designed for Information-Centric Networking (ICN). Through a structured classification framework, the work comprehensively reviews and compares existing NDN caching strategies in terms of their operational mechanisms, performance metrics, strengths, and limitations, thereby clarifying the appropriate application scenarios for each approach. The analysis not only establishes a theoretical foundation for designing high-performance caching mechanisms but also identifies promising directions for future research, offering clear academic guidance to the field.
This paper identifies a critical gap in high-performance data transfer research: an overemphasis on network bandwidth while neglecting end-to-end bottlenecks—including latency, TCP congestion control, host CPU limitations, and virtualization—leading to severe discrepancies between benchmark results and real-world production performance. To address this, the authors propose a hardware–software co-design paradigm and develop a latency-programmable testbed. Leveraging high-fidelity wide-area network (WAN) modeling and cross-continental 100 Gbps measurements (Switzerland–California), they systematically isolate key constraints at the network edge and host side. Results demonstrate that primary bottlenecks reside predominantly at the network edge—not the core—and that stable, predictable data movement is achieved across 1–100+ Gbps. This significantly enhances performance fidelity in complex, heterogeneous environments.
This study addresses the unpredictable end-to-end latency in cloud virtualized environments, which stems from virtualization overheads in CPU, I/O, and network resources. Through systematic network measurement experiments across diverse virtualization platforms—including KVM, LXC, and Docker—under multidimensional workload conditions, the authors collect packet round-trip time data to construct a high-quality dataset suitable for machine learning–based network performance modeling. By integrating data preprocessing, correlation analysis, dimensionality reduction, and clustering techniques, this work presents the first quantitative evaluation of latency impacts across multiple virtualization technologies. The resulting dataset effectively supports network performance prediction and intelligent resource scheduling, providing an empirical foundation for performance optimization in cloud environments.
To address the challenge in distributed storage systems of simultaneously achieving low storage overhead, high reliability, and low repair traffic—where replication and erasure coding (EC) individually fall short—this paper proposes HyRES, a network-scale-aware hybrid storage scheme. HyRES innovatively unifies replication and EC within a single coherent framework, rather than merely combining them. It introduces a dynamic tiered encoding strategy and a scale-adaptive repair scheduling mechanism to jointly optimize storage cost, file loss probability (FLP), and cross-network repair traffic. Theoretical modeling and large-scale simulations demonstrate that, under identical fault tolerance guarantees, HyRES reduces storage overhead by approximately 40% compared to pure replication, lowers FLP by over 50% relative to conventional EC, and significantly mitigates the scaling of repair traffic with increasing network size.
This work addresses the performance bottleneck in traditional databases caused by reliance on the kernel TCP stack, which becomes a CPU-intensive black box in high-speed cloud networks. To overcome this limitation, the authors propose a dual-channel networking paradigm that decouples database communication into a high-performance data channel based on user-space UDP and a reliable control channel leveraging kernel TCP. By co-designing the database with modern NIC hardware features, this approach simultaneously achieves low latency, high throughput, and strong reliability. Empirical results demonstrate that the system saturates a 200 Gbit/s network link using only three CPU cores during distributed shuffle operations and enables a replicated key-value store to process millions of messages per second.