🤖 AI Summary
This study addresses the unclear relationships among information sources, generation strategies, and quality evidence in test case generation using large language models (LLMs). Through a systematic literature review of 95 studies, this work constructs a multidimensional taxonomy and a benchmark analysis framework. Specifically, it proposes a four-dimensional classification system that elucidates how execution feedback influences oracle independence. Furthermore, it establishes a unified theoretical framework connecting the generation process with quality assessment. By identifying independent oracle evaluation as a critical yet underexplored dimension, this research formulates a future agenda centered on rigorous, oracle-independent quality measurement for LLM-generated test cases.
📝 Abstract
Large language models are increasingly used to generate test scenarios, executable test suites, assertions, interaction sequences, and fuzzing infrastructure. These artifacts serve different purposes and rely on different sources of information about correct behavior. A test that increases implementation coverage, a test that agrees with a reference program, and a test that detects a requirements violation therefore provide distinct kinds of evidence. We present a structured narrative survey that connects the generation process to the evidence used to justify test quality. Drawing on 95 curated source records through 29 September 2026, we organize the field along four dimensions: testing objectives and artifacts, information available during training and generation, generation and learning mechanisms, and evaluation evidence. We synthesize work on unit and requirements-based testing, test oracles, API and GUI testing, fuzzing, and learned test generators. The resulting framework explains how execution feedback can improve executability while also affecting the independence of an oracle, why specification provenance and training-time reference supervision matter for comparisons, and how downstream code-selection gains differ from test correctness. We use these distinctions to organize benchmarks and derive a research agenda for independent oracle assessment, budget-aware generation, repository-scale evaluation, and test maintenance.