Score
Designs and executes measurements and analyses that quantify how available memory or storage capacity affects system performance, producing tradeoff curves and identifying optimal, fragile, or regime-dependent capacities. Builds evaluation protocols, metrics, and models (for example information‑bottleneck or coordination-across-channel analyses) to compare architectures and predict how changes in capacity alter behavior.
AI infrastructure confronts multidimensional physical and economic constraints—including power, thermal management, water usage, interconnect bandwidth, memory capacity, and data throughput—while existing metrics (e.g., PUE, TCO) are siloed and fail to capture the coupled trade-offs among energy efficiency, performance, and cost, hindering cross-layer co-optimization. To address this, we propose a unified measurement architecture grounded in a 6×3 cross-layer taxonomy—spanning facility, network, compute, storage, software, and application layers, each annotated with physical, computational, and economic semantics—and introduce the Measurement Propagation Graph (MPG) to enable, for the first time, system-level, three-dimensional relational modeling. Leveraging systematic literature review, meta-analysis, and graph-based modeling, our framework integrates heterogeneous, multi-source metrics. It supports benchmarking, capacity planning, and total cost of ownership analysis, substantially enhancing interpretability of AI cluster efficiency frontiers and enabling rigorous multi-objective optimization.
To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.
This work addresses the challenges of SLO violations and resource inefficiency in machine learning model serving caused by inadequate capacity planning. To this end, the authors propose an adaptive, feedback-driven load testing framework that formalizes the ML serving load testing process for the first time. The framework incorporates real-traffic-based workload calibration and a warm-up mechanism, combined with adaptive search, performance signal feedback control, convergence detection, and GPU monitoring to efficiently estimate the maximum sustainable throughput under SLO constraints. Evaluation across 14 industrial cases demonstrates that the approach reduces capacity estimation error from approximately 30% to 2–6%, with the warm-up mechanism improving accuracy by 22.2%. This significantly mitigates deployment incidents and enhances GPU resource utilization efficiency.
Modeling complex inter-parameter dependencies in configurable software systems remains challenging for performance tuning. Method: This paper pioneers a fitness landscape (FL) perspective to reconstruct performance analysis, modeling the high-dimensional configuration space as a structured terrain—departing from conventional isolated-point evaluation. It integrates graph-based data mining, fitness landscape analysis (FLA), and large-scale sampling (86 million configurations) across three real-world systems and 32 workload types. Contribution/Results: The study uncovers six universal landscape patterns, enabling robust identification of local optima and precise topological characterization of configuration terrain. These findings substantially deepen insights into black-box system performance, providing both a novel theoretical foundation and reusable practical guidelines for automated configuration tuning and performance modeling.
Configuration space explosion complicates performance impact modeling, while gray-box approaches rely on structural knowledge (e.g., module execution graphs) to improve model accuracy—yet the mechanisms by which structural features (e.g., number of modules or configuration options) and structural knowledge influence modeling difficulty and optimization potential remain unclear. Method: We formally define “modeling hardness” and “improvement opportunity,” establishing an analytical framework and matrix to quantify the interplay among system structural complexity, structural knowledge level, and modeling benefit. Controlled experiments on synthetic systems integrate module execution graph analysis with gray-box modeling. Contribution/Results: We identify module count and configuration option count as dominant determinants of modeling hardness. Under high hardness, strong structural knowledge significantly increases improvement opportunity. Structural knowledge primarily enhances ranking accuracy, whereas hardness predominantly degrades prediction accuracy. Our findings provide theoretical foundations and strategic guidance for allocating structural knowledge investment according to specific modeling objectives.
This work addresses the significant discrepancies between existing memory simulators and real hardware when predicting the performance of advanced memory systems, compounded by a lack of reliable validation methodologies. To tackle this issue, we propose the first multi-perspective co-validation framework that systematically evaluates simulation accuracy from three complementary dimensions: the memory simulator itself, the CPU–memory interface, and application-level behavior. Our analysis reveals that inaccuracies at the interface layer are a primary source of simulation distortion. Building on this insight, we integrate mainstream simulators—Ramulator, Ramulator2, and DRAMsim3—into the ZSim platform and implement targeted corrections and enhancements at the interface layer. Experimental results demonstrate that the refined simulators achieve substantially improved fidelity across diverse workloads, yielding predictions that closely align with real-system performance.
This work addresses the computational and memory bottlenecks that hinder efficient scaling in large model training. To overcome the limitations of conventional point-wise optimizations, the authors propose a throughput-centric strategy that systematically integrates multiple techniques: optimized data loading (OVERLORD), CPU memory offloading (DeepSpeed ZeRO-Offload), distributed compilation (Triton-distributed), and hardware-level dynamic voltage and frequency scaling (DVFS). This holistic approach achieves a 4.5% improvement in end-to-end training throughput, substantially reduces training costs, and enables efficient training of models significantly larger than the memory capacity of a single GPU.
This work addresses the challenge that state-dependent behaviors in modern computing environments—such as those introduced by adaptive system mechanisms—induce time-dependent biases in traditional software benchmarking, undermining reliable performance comparisons. The paper reframes benchmarking as a decision problem aimed at identifying the fastest program and introduces an experimental paradigm centered on pairwise performance comparisons, thereby avoiding strong assumptions about modeling system dynamics. By leveraging contrastive estimators, consistent statistical inference, and test strategies under finite evaluation budgets, the approach eliminates program-specific biases without relying on absolute performance metrics. The method provides asymptotic guarantees for correct decisions in stateful environments, offering a robust and reliable benchmarking framework for performance-sensitive software development.
This work addresses the inefficiency of traditional approaches to predicting workload performance under varying memory configurations, which typically rely on time-consuming simulations or repeated measurements. The study reveals, for the first time, a predictable relationship between cycles per instruction (CPI) and maximum memory stall across diverse workloads. By leveraging hardware performance counters collected from a single native execution—combined with mechanistic insights and empirical data—the authors construct a regression model that enables highly accurate, simulation-free first-order performance prediction. Evaluated across six machine configurations and two simulators, the method reduces CPI prediction error by 2× compared to the best existing single-run techniques. On ARM servers, it achieves a median error of 12.7% and a 90th-percentile error of 35.9%, maintaining robust accuracy even when extrapolating to memory latencies up to 8× higher than baseline.