🤖 AI Summary
To address the I/O bottleneck in ML training—causing low GPU utilization (often <50%)—this paper proposes a data-driven approach for I/O performance prediction and storage configuration optimization. We conduct systematic benchmarking across 141 configurations spanning diverse storage backends (NVMe SSDs, network-attached storage, in-memory filesystems), data formats, and access patterns. Leveraging key features—including batch size and throughput—we train an XGBoost regression model achieving an R² of 0.991 and a mean absolute error of only 11.8%. The model enables minute-scale recommendation of optimal storage configurations, accelerating configuration search by over an order of magnitude compared to empirical trial-and-error. Our core contribution is the first general-purpose, ML-training-pipeline-aware I/O performance prediction framework, which significantly improves GPU utilization and end-to-end training efficiency. All code and benchmark data are fully open-sourced, ensuring strong reproducibility and extensibility.
📝 Abstract
Modern machine learning training is increasingly bottlenecked by data I/O rather than compute. GPUs often sit idle at below 50% utilization waiting for data. This paper presents a machine learning approach to predict I/O performance and recommend optimal storage configurations for ML training pipelines. We collected 141 observations through systematic benchmarking across different storage backends (NVMe SSD, network-attached storage, in-memory filesystems), data formats, and access patterns, covering both low-level I/O operations and full training pipelines. After evaluating seven regression models and three classification approaches, XGBoost achieved the best performance with R-squared of 0.991, predicting I/O throughput within 11.8% error on average. Feature importance analysis revealed that throughput metrics and batch size are the primary performance drivers. This data-driven approach can reduce configuration time from days of trial-and-error to minutes of predictive recommendation. The methodology is reproducible and extensible to other resource management problems in ML systems. Code and data are available at https://github.com/knkarthik01/gpu_storage_ml_project