🤖 AI Summary
Boundary value analysis (BVA) for complex software systems faces challenges in generating effective test data, particularly for high-dimensional, nonlinear, or rare boundary conditions, leading to inadequate coverage. Method: This study presents the first systematic empirical evaluation of large language models (LLMs) in white-box BVA. We propose a prompt-engineering–based, LLM-driven test generation method and conduct rigorous empirical assessment using fault injection and multi-dimensional coverage metrics (e.g., statement and branch coverage). Results: LLM-generated test suites match traditional approaches in detecting common boundary defects, achieving a 12% average improvement in fault detection rate. However, they exhibit limited effectiveness in high-dimensional, nonlinear, and infrequent boundary scenarios—yielding 18% lower statement and branch coverage. Contribution: This work establishes the first BVA-specific empirical benchmark for LLMs in software testing and reveals their practical effectiveness boundaries, offering foundational methodological insights for LLM-augmented testing.
📝 Abstract
As software systems grow more complex, automated testing has become essential to ensuring reliability and performance. Traditional methods for boundary value test input generation can be time-consuming and may struggle to address all potential error cases effectively, especially in systems with intricate or highly variable boundaries. This paper presents a framework for assessing the effectiveness of large language models (LLMs) in generating boundary value test inputs for white-box software testing by examining their potential through prompt engineering. Specifically, we evaluate the effectiveness of LLM-based test input generation by analyzing fault detection rates and test coverage, comparing these LLM-generated test sets with those produced using traditional boundary value analysis methods. Our analysis shows the strengths and limitations of LLMs in boundary value generation, particularly in detecting common boundary-related issues. However, they still face challenges in certain areas, especially when handling complex or less common test inputs. This research provides insights into the role of LLMs in boundary value testing, underscoring both their potential and areas for improvement in automated testing methods.