🤖 AI Summary
This study addresses the limitation of learning-based mutant selection, which tends to over-concentrate testing budgets on specific code locations, resulting in blind spots and missed fault detection. To mitigate this, we propose LinePool, a model-agnostic stratified sampling method that allocates the budget hierarchically across source code lines. This approach preserves ranking signals while balancing spatial coverage with behavioral diversity. Experiments across multiple datasets demonstrate that LinePool significantly improves spatial coverage and reduces kill set overlap. Furthermore, it effectively enhances fault revelation rates under strict budget constraints and exhibits superior robustness to ranking signal shifts compared to existing clustering and diversification methods.
📝 Abstract
Learning-based mutant selection is used to reduce the cost of mutation testing by ranking and selecting the mutants that seem to be the most promising. Prior work has shown great potential for this approach in fault revelation and subsuming-mutant selection. However, the risks and benefits of these approaches -- especially the distribution of selected mutants across code locations and how that distribution affects behavioral diversity and fault revelation -- remain unclear. We investigate this distribution and its effects, and introduce LinePool, a simple, model-agnostic stratification step that distributes selections across source lines while retaining the underlying ranking signal. Across two datasets and budgets (2% and 5% of the mutants), we find that learning-based selection concentrates the budget on a few code locations, and may thereby systematically introduce'blind spots', untested code areas, allowing faults to escape detection. On larger programs, the selected mutants also exhibit substantial overlap in kill behavior. LinePool substantially increases spatial coverage and reduces kill-set overlap. It also improves fault revelation at tight budgets and makes selection more robust to imperfect or shifted ranking signals. Comparisons with clustering and established diversification methods show that none consistently outperforms LinePool across datasets, supporting its use as a simple diversification step. In general, our results suggest that behavioral diversity should be considered alongside effectiveness when designing and evaluating mutant-selection methods, and that exploiting structural elements of the program is one simple way to obtain such diversity while maintaining effectiveness.