Scaling Performance of Large Language Model Pretraining

📅 2025-09-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large language model (LLM) pretraining faces significant challenges, including prohibitive computational costs, opaque scaling laws, and a lack of practical guidance for large-scale distributed training. Method: This work systematically investigates the performance scaling mechanisms of LLM pretraining pipelines at the hundred-node scale, focusing on three key directions: optimization of distributed training architecture, cross-node efficient dataset management, and deep scaling of data parallelism—aiming to maximize GPU resource utilization. Through empirical analysis, we quantify the interplay among communication overhead, I/O bottlenecks, and parallelism degree, and establish a reproducible framework for large-scale training performance tuning. Contribution/Results: Our study bridges a critical gap in the public literature on engineering practices for ultra-large-scale LLM training, delivering a practical, deployable technical pathway and actionable guidelines for efficient pretraining on thousand-GPU clusters.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsData Mining & Knowledge Management: Scalability, Parallel & Distributed Systems

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Large language models (LLMs) show best-in-class performance across a wide range of natural language processing applications. Training these models is an extremely computationally expensive task; frontier Artificial Intelligence (AI) research companies are investing billions of dollars into supercomputing infrastructure to train progressively larger models on increasingly massive datasets. Unfortunately, information about the scaling performance and training considerations of these large training pipelines is scarce in public literature. Working with large-scale datasets and models can be complex and practical recommendations are scarce in the public literature for tuning training performance when scaling up large language models. In this paper, we aim to demystify the large language model pretraining pipeline somewhat - in particular with respect to distributed training, managing large datasets across hundreds of nodes, and scaling up data parallelism with an emphasis on fully leveraging available GPU compute capacity.
Problem

Research questions and friction points this paper is trying to address.

Optimizing distributed training for large language models
Managing massive datasets across hundreds of nodes
Maximizing GPU utilization during LLM pretraining scaling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distributed training across hundreds of nodes
Scaling data parallelism for GPU utilization
Managing large datasets in training pipelines
🔎 Similar Papers
A
Alexander Interrante-Grant
Lincoln Laboratory, Massachusetts Institute of Technology
C
Carla Varela-Rosa
Lincoln Laboratory, Massachusetts Institute of Technology
S
Suhaas Narayan
Lincoln Laboratory, Massachusetts Institute of Technology
C
Chris Connelly
Lincoln Laboratory, Massachusetts Institute of Technology
A
Albert Reuther
Lincoln Laboratory, Massachusetts Institute of Technology