🤖 AI Summary
This study addresses the lack of systematic guidance for selecting evaluation metrics in combinatorial interaction testing by presenting the first comprehensive survey of black-box static metrics. Through literature analysis and large-scale software testing experiments, it empirically validates the correlation between various metrics and fault detection effectiveness. The research confirms that value combination coverage serves as an effective predictor of fault detection, reveals the cost-efficiency advantages of distribution-based metrics, and delivers practical guidelines for metric selection. By filling the gap in systematic empirical research within this field, this work provides practitioners with a scientifically grounded basis for evaluating test suites.
📝 Abstract
Combinatorial interaction testing (CIT) is a black-box testing method that has received extensive attention in both research and practice over recent years. Its primary objective is to construct an effective combinatorial test suite that detects software failures caused by parameter interactions. As a fundamental component of the CIT testing process, the evaluation metric plays a critical role in assessing and comparing combinatorial test suites, as well as in evaluating various test generation techniques. For CIT practitioners, selecting an appropriate evaluation metric is both important and challenging, given the wide variety of available options. Nevertheless, no prior work has systematically addressed this problem. To fill this gap, this paper first provides a comprehensive survey of black-box evaluation metrics for combinatorial test suites, offering rigorous definitions, clear classifications, illustrative examples, and complexity analyses. We then conduct an extensive empirical study involving eight open-source projects, encompassing 32 test scenarios and 295,624 combinatorial test suites. In this study, we examine the correlation between each static evaluation metric and fault-detection effectiveness using two correlation measures. Experimental results show that the Value Combination Coverage (VCC) metric serves as a valid predictor for test-suite evaluation. However, distribution-based metrics generally incur lower computational costs than interaction coverage-based ones. The choice of an appropriate metric should also account for the test suite's inherent properties, as these characteristics can substantially influence the effectiveness of the evaluation. Finally, we provide practical guidelines to assist CIT practitioners in selecting suitable evaluation metrics for assessing or comparing combinatorial test suites.