🤖 AI Summary
Existing evaluations of LLM-generated code tests often overestimate test quality by relying on single reference solutions and overlooking alternative implementations. To address this limitation, this work constructs a benchmark comprising 300 tasks with 3,000 diverse implementations and introduces the first joint success function metric based on multi-candidate validation. Furthermore, we propose TestHelix, a framework that integrates heterogeneous synthesis, peer cross-validation, and recursive self-improvement mechanisms to optimize test generation. Experimental results demonstrate that baseline methods achieve a joint success rate of only 28%, whereas TestHelix yields an improvement of approximately 9 percentage points over native configurations, significantly enhancing the robustness of code testing.
📝 Abstract
Large language model (LLM) coding agents have advanced test generation across diverse programming tasks. However, the common practice of evaluating tests against a single reference solution overlooks alternative valid implementations and can overstate test quality. We introduce TestPrism, comprising 300 test tasks from 17 sources and 3000 candidate implementations, evenly split between valid and invalid solutions. Its primary metric, Joint Success Function, requires the generated tests to fail on the initial program state, accept every valid candidate, and reject every invalid candidate. Across fourteen baseline coding agent configurations, Joint Success Function reaches only 28.00%, whereas single reference success reaches 59.67%. Our analysis reveals missed behaviors, unsupported assertions, and faulty test construction. To address these weaknesses, we introduce TestHelix, which combines heterogeneous synthesis of test and repair pairs with peer cross validation and recursive self improvement (RSI). Across two models, TestHelix improves Joint Success Function by 8.67 to 9.00 percentage points over the native harness comparators in the TestHelix evaluation