WitnessGym: Benchmarking Coding Agents on the Construction of Bug Witnesses

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of high-quality automated benchmarks for evaluating coding agents in bug verification by proposing an automated benchmark construction framework based on bug injection and behavior-preserving transformations. The approach injects defects into real-world Java projects while retaining executable witnesses, integrating dynamic execution adapters with structure-preserving transformation algorithms to enable automated verification via authentic test suites. This methodology overcomes the limitations of reusing historical bugs or relying on manually constructed data. Accordingly, a highly realistic benchmark comprising 1,300 cases is established. Experimental results demonstrate that mainstream models struggle to effectively generate executable witnesses even when defect patterns are known, revealing a fundamental technical bottleneck in current coding agents.
📝 Abstract
Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution. Such evidence makes audit findings actionable, yet benchmark evaluation is difficult when cases reuse public historical bugs and witnesses or require manual construction. We present WitnessGym, an automated framework for constructing bug-validation benchmarks through bug injection. It injects bugs into test-reached paths of real projects, rebuilds each project, and retains cases exposed by a construction-time witness. Bug specifications and execution adapters allow extension to additional bug types and languages. Bug-preserving transformations vary the surrounding structure while preserving the witness behavior. Based on real-world Java projects with test suites, WitnessGym automatically constructs 1,300 benchmark cases. The injected patches resemble historical bug patches and are difficult for the two evaluated models to distinguish in blinded comparisons. We evaluate four coding agent frameworks in six framework/model pairings across bug types, execution contexts, and transformation depths. Witness construction remains difficult even when the bug pattern is known. Our framework, benchmark cases, and evaluation scripts are available.
Problem

Research questions and friction points this paper is trying to address.

Bug Validation
Coding Agents
Benchmark Construction
Bug Witness
Automated Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bug Injection
Automated Benchmarking
Coding Agents
Bug Witness
Bug-preserving Transformation