Boundary Value Test Input Generation Using Prompt Engineering with LLMs: Fault Detection and Coverage Analysis

📅 2025-01-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Boundary value analysis (BVA) for complex software systems faces challenges in generating effective test data, particularly for high-dimensional, nonlinear, or rare boundary conditions, leading to inadequate coverage. Method: This study presents the first systematic empirical evaluation of large language models (LLMs) in white-box BVA. We propose a prompt-engineering–based, LLM-driven test generation method and conduct rigorous empirical assessment using fault injection and multi-dimensional coverage metrics (e.g., statement and branch coverage). Results: LLM-generated test suites match traditional approaches in detecting common boundary defects, achieving a 12% average improvement in fault detection rate. However, they exhibit limited effectiveness in high-dimensional, nonlinear, and infrequent boundary scenarios—yielding 18% lower statement and branch coverage. Contribution: This work establishes the first BVA-specific empirical benchmark for LLMs in software testing and reveals their practical effectiveness boundaries, offering foundational methodological insights for LLM-augmented testing.

Technology Category

Machine Learning: Evaluation and AnalysisConstraint Satisfaction and Optimization: Satisfiability Modulo TheoriesNatural Language Processing: Generation

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Web evaluation methodologies and metrics
📝 Abstract
As software systems grow more complex, automated testing has become essential to ensuring reliability and performance. Traditional methods for boundary value test input generation can be time-consuming and may struggle to address all potential error cases effectively, especially in systems with intricate or highly variable boundaries. This paper presents a framework for assessing the effectiveness of large language models (LLMs) in generating boundary value test inputs for white-box software testing by examining their potential through prompt engineering. Specifically, we evaluate the effectiveness of LLM-based test input generation by analyzing fault detection rates and test coverage, comparing these LLM-generated test sets with those produced using traditional boundary value analysis methods. Our analysis shows the strengths and limitations of LLMs in boundary value generation, particularly in detecting common boundary-related issues. However, they still face challenges in certain areas, especially when handling complex or less common test inputs. This research provides insights into the role of LLMs in boundary value testing, underscoring both their potential and areas for improvement in automated testing methods.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Software Testing
Test Coverage
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Boundary Test Data Generation
Enhanced White-box Testing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xiujing Guo
Graduate School of Information Science and Technology, Osaka University, Osaka, Japan
C
Chen Li
Graduate School of Informatics, Nagoya University, Nagoya, Japan
Tatsuhiro Tsuchiya
Tatsuhiro Tsuchiya
Professor, Osaka University
Software TestingVerificationFault Tolerance