Predictive Modeling of I/O Performance for Machine Learning Training Pipelines: A Data-Driven Approach to Storage Optimization

📅 2025-12-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the I/O bottleneck in ML training—causing low GPU utilization (often <50%)—this paper proposes a data-driven approach for I/O performance prediction and storage configuration optimization. We conduct systematic benchmarking across 141 configurations spanning diverse storage backends (NVMe SSDs, network-attached storage, in-memory filesystems), data formats, and access patterns. Leveraging key features—including batch size and throughput—we train an XGBoost regression model achieving an R² of 0.991 and a mean absolute error of only 11.8%. The model enables minute-scale recommendation of optimal storage configurations, accelerating configuration search by over an order of magnitude compared to empirical trial-and-error. Our core contribution is the first general-purpose, ML-training-pipeline-aware I/O performance prediction framework, which significantly improves GPU utilization and end-to-end training efficiency. All code and benchmark data are fully open-sourced, ensuring strong reproducibility and extensibility.

Technology Category

Machine Learning: Hardware-aware MLSearch and Optimization: Learning to SearchComputer Vision: Learning & Optimization for CV

Application Category

User Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendationSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Modern machine learning training is increasingly bottlenecked by data I/O rather than compute. GPUs often sit idle at below 50% utilization waiting for data. This paper presents a machine learning approach to predict I/O performance and recommend optimal storage configurations for ML training pipelines. We collected 141 observations through systematic benchmarking across different storage backends (NVMe SSD, network-attached storage, in-memory filesystems), data formats, and access patterns, covering both low-level I/O operations and full training pipelines. After evaluating seven regression models and three classification approaches, XGBoost achieved the best performance with R-squared of 0.991, predicting I/O throughput within 11.8% error on average. Feature importance analysis revealed that throughput metrics and batch size are the primary performance drivers. This data-driven approach can reduce configuration time from days of trial-and-error to minutes of predictive recommendation. The methodology is reproducible and extensible to other resource management problems in ML systems. Code and data are available at https://github.com/knkarthik01/gpu_storage_ml_project
Problem

Research questions and friction points this paper is trying to address.

Predicts I/O performance for ML training pipelines
Recommends optimal storage configurations to reduce idle time
Uses data-driven modeling to replace trial-and-error with predictions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses XGBoost to predict I/O performance accurately
Analyzes storage configurations via systematic data collection
Reduces configuration time from days to minutes
K
Karthik Prabhakar
University of Texas at Austin