CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of benchmarks evaluating agents' long-horizon capabilities in developing static analysis checkers from scratch. We propose the first end-to-end evaluation framework for checker synthesis, constructing an executable benchmark comprising 300 multilingual CVE tasks. Methodologically, our approach integrates compilation feedback loops, patch localization, and tool-use tracing, quantifying diagnostic comparisons and false positive rates through independently reconstructed artifacts. Experiments across 21 configurations reveal an average Pass@1 of only 32.3% (peaking at 45.33%), demonstrating that current agents remain unable to reliably complete this complex, end-to-end workflow.
📝 Abstract
Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing coding-agent benchmarks focus on tasks such as patch generation or vulnerability detection and rarely assess whether an agent can develop a working checker in a repository from start to finish. We introduce CheckerBench, an executable benchmark of 300 tasks derived from 297 CVEs across 167 repositories, 85 CWEs, and five language ecosystems. Each task includes vulnerable and fixed revisions, a pinned analysis environment, and a checker scaffold. We further introduce CheckerLab, a common evaluation framework that independently rebuilds submitted checkers and measures vulnerable-fixed diagnostic contrast, patch localization, false positives, and tool use. Across 21 model-harness configurations and three independent repeats per configuration, mean Pass@1 is 32.30%, while the best reaches 45.33%. These results show that reliable, reusable checker development remains challenging for current coding agents.
Problem

Research questions and friction points this paper is trying to address.

Static-Analysis Checker Synthesis
Coding Agents
Benchmark Evaluation
Long-Horizon Tasks
Vulnerability Detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Static-Analysis Checker Synthesis
Executable Benchmark
Coding Agents
Evaluation Framework
Long-Horizon Agents
🔎 Similar Papers