🤖 AI Summary
Conventional fuzz testing evaluation primarily focuses on holistic performance metrics, overlooking how fine-grained program characteristics influence fuzzer effectiveness.
Method: We introduce the first program-feature-oriented fuzzing benchmark, generating 153 controllable C programs within a synthetic framework using 10 configurable parameters. For the first time, seven critical structural features—including branch depth, loop nesting level, and memory reference patterns—are explicitly modeled as tunable dimensions, overcoming the intrinsic-structure blindness of existing black-box and gray-box benchmarks.
Contribution/Results: Leveraging an automated evaluation framework, we systematically assess 11 mainstream fuzzers (e.g., AFL++ and LibFuzzer). Our analysis reveals strong sensitivity of code coverage to program features: for instance, high loop nesting reduces coverage by up to 63% for certain tools. These findings empirically validate the necessity and efficacy of feature-aware fuzzing evaluation.
📝 Abstract
Fuzzing is a powerful software testing technique renowned for its effectiveness in identifying software vulnerabilities. Traditional fuzzing evaluations typically focus on overall fuzzer performance across a set of target programs, yet few benchmarks consider how fine-grained program features influence fuzzing effectiveness. To bridge this gap, we introduce a novel benchmark designed to generate programs with configurable, fine-grained program features to enhance fuzzing evaluations. We reviewed 25 recent grey-box fuzzing studies, extracting 7 program features related to control-flow and data-flow that can impact fuzzer performance. Using these features, we generated a benchmark consisting of 153 programs controlled by 10 fine-grained configurable parameters. We evaluated 11 popular fuzzers using this benchmark. The results indicate that fuzzer performance varies significantly based on the program features and their strengths, highlighting the importance of incorporating program characteristics into fuzzing evaluations.