🤖 AI Summary
This study addresses the inefficiency bottleneck of GPU-intensive continuous integration (CI) in large language model (LLM) training by proposing an optimization framework that integrates dynamic test selection with risk-prioritized scheduling. The approach leverages runtime evidence to precisely identify affected tests, prunes contextually equivalent test cases, and synergizes high-dimensional workload optimization to prioritize high-risk tests, thereby transcending the conventional full regression testing paradigm. Deployed at scale within ByteDance, the proposed framework reduces CI latency by 77.5% and GPU resource consumption by 63.9%, while simultaneously improving code coverage by 3.2%. These results demonstrate a significant advancement in industrial-grade LLM CI efficiency, offering a practical solution for scalable and cost-effective model development pipelines.
📝 Abstract
As large language models (LLMs) keep growing in size and complexity, their training frameworks evolve at a rapid pace as well. Therefore, continuous integration (CI) is critical for maintaining the quality and stability of these frameworks. However, unlike traditional software, CI for LLM training frameworks relies on GPU-intensive tests, which usually involve complete model training or evaluation. This leads CI itself to become a new bottleneck for fast-paced development. In this paper, we introduce FastCI, a framework that improves the efficiency of CI for LLM training frameworks. FastCI leverages runtime evidence to select affected tests and prune tests that execute changed code in equivalent contexts. Then FastCI prioritizes high-risk tests to expose potential failures earlier, and optimizes test workloads along dimensions outside the intended validation scope of each test. Evaluated on the CI workload of our LLM training framework, FastCI reduces the CI latency by 77.5% and the GPU resource usage by 63.9%, while improving the modified code coverage retention by 3.2%, compared with the currently deployed CI pipelines. FastCI has now been integrated into the CI pipelines of our LLM training framework at ByteDance.