Hiding Latencies in Network-Based Image Loading for Deep Learning

πŸ“… 2025-03-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
To address data loading bottlenecks and local storage overload caused by high network/storage latency in deep learning image training, this paper proposes a unified data management framework built upon scalable NoSQL databases (e.g., Cassandra). The method introduces three key innovations: (1) an adaptive out-of-order incremental prefetching mechanism that effectively masks I/O latency across intercontinental high-latency networks; (2) the first low-overhead integration of scalable NoSQL backends with native PyTorch/TensorFlow data loaders; and (3) RDMA- and HTTP/3–enabled transport optimizations. Experimental evaluation demonstrates that the framework achieves 92% of local SSD throughput, reduces metadata query latency by three orders of magnitude, and decreases per-node storage pressure by 76%. These results validate its effectiveness in decoupling compute scalability from local storage constraints while maintaining high training efficiency.

Technology Category

Search and Optimization: Distributed SearchComputer Vision: Learning & Optimization for CVMachine Learning: Learning on the Edge & Model Compression

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Data management and stream processing for Web, mobile and wireless applicationsSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesGraph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphs
πŸ“ Abstract
In the last decades, the computational power of GPUs has grown exponentially, allowing current deep learning (DL) applications to handle increasingly large amounts of data at a progressively higher throughput. However, network and storage latencies cannot decrease at a similar pace due to physical constraints, leading to data stalls, and creating a bottleneck for DL tasks. Additionally, managing vast quantities of data and their associated metadata has proven challenging, hampering and slowing the productivity of data scientists. Moreover, existing data loaders have limited network support, necessitating, for maximum performance, that data be stored on local filesystems close to the GPUs, overloading the storage of computing nodes. In this paper we propose a strategy, aimed at DL image applications, to address these challenges by: storing data and metadata in fast, scalable NoSQL databases; connecting the databases to state-of-the-art loaders for DL frameworks; enabling high-throughput data loading over high-latency networks through our out-of-order, incremental prefetching techniques. To evaluate our approach, we showcase our implementation and assess its data loading capabilities through local, medium and high-latency (intercontinental) experiments.
Problem

Research questions and friction points this paper is trying to address.

Reducing network and storage latencies in deep learning image loading
Managing large data and metadata efficiently for data scientists
Enhancing data loader performance over high-latency networks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Using NoSQL databases for data and metadata
Connecting databases to advanced DL loaders
Prefetching data out-of-order to reduce latency
πŸ”Ž Similar Papers
No similar papers found.