🤖 AI Summary
This study addresses the challenge that evidence of compiler-introduced security vulnerabilities is often fragmented and difficult to audit by introducing the first auditable Source-IR paired dataset. Methodologically, candidate samples are screened through a combination of data mining and LLVM IR static analysis, followed by double-blind manual annotation and inter-annotator agreement checks to rigorously distinguish genuine vulnerabilities from hard negative examples. The resulting dataset comprises 429 GCC/LLVM cases, including 280 confirmed vulnerabilities and 149 hard negatives, each accompanied by source code, intermediate representations, and multidimensional labels to support multimodal analysis. This work validates the feasibility of joint source-IR analysis and establishes a standardized benchmark for future research in compiler security.
📝 Abstract
Compiler-introduced security bugs (CISBs) arise when an optimization, lowering, or instrumentation decision changes a security-relevant property of the generated program. They are difficult to study because their evidence is distributed across issue reports, reduced tests, historical configurations, and compiler artifacts; a security-related report also does not imply that every associated reduction establishes a security-bearing compiler failure. We present CISB-Bench, an auditable dataset of 429 exact C-program rows mined from GCC and LLVM. Each row contains its C reduction, a standardized LLVM IR analysis bundle at -O0 through -O3, public provenance, a final binary label, and a primary mechanism or boundary annotation. Two reviewers independently labeled the fixed corpus, agreeing on 369 rows (86.0%, Cohen's kappa=0.662); the 60 disagreements were adjudicated. The final dataset comprises 280 CISBs and 149 hard non-CISB cases. The prediction task is to recover this reviewed exact-row label from the supplied artifacts; it is not a claim that standardized IR alone reproduces every historical compiler failure. We characterize the security mechanisms and evidence boundaries represented by the corpus, and demonstrate how its paired artifacts support source-only, IR-aware, and joint analyses. CISB-Bench provides a reusable, inspectable target for compiler-security mining and detection research.