TestPrism: Rethinking Test Evaluation Beyond a Single Reference

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing evaluations of LLM-generated code tests often overestimate test quality by relying on single reference solutions and overlooking alternative implementations. To address this limitation, this work constructs a benchmark comprising 300 tasks with 3,000 diverse implementations and introduces the first joint success function metric based on multi-candidate validation. Furthermore, we propose TestHelix, a framework that integrates heterogeneous synthesis, peer cross-validation, and recursive self-improvement mechanisms to optimize test generation. Experimental results demonstrate that baseline methods achieve a joint success rate of only 28%, whereas TestHelix yields an improvement of approximately 9 percentage points over native configurations, significantly enhancing the robustness of code testing.
📝 Abstract
Large language model (LLM) coding agents have advanced test generation across diverse programming tasks. However, the common practice of evaluating tests against a single reference solution overlooks alternative valid implementations and can overstate test quality. We introduce TestPrism, comprising 300 test tasks from 17 sources and 3000 candidate implementations, evenly split between valid and invalid solutions. Its primary metric, Joint Success Function, requires the generated tests to fail on the initial program state, accept every valid candidate, and reject every invalid candidate. Across fourteen baseline coding agent configurations, Joint Success Function reaches only 28.00%, whereas single reference success reaches 59.67%. Our analysis reveals missed behaviors, unsupported assertions, and faulty test construction. To address these weaknesses, we introduce TestHelix, which combines heterogeneous synthesis of test and repair pairs with peer cross validation and recursive self improvement (RSI). Across two models, TestHelix improves Joint Success Function by 8.67 to 9.00 percentage points over the native harness comparators in the TestHelix evaluation
Problem

Research questions and friction points this paper is trying to address.

test evaluation
large language models
coding agents
single reference solution
test generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

TestPrism
Joint Success Function
TestHelix
Recursive Self Improvement
Heterogeneous Synthesis
💼 Related Jobs
No related jobs found.
H
Han Li
Nanjing University
L
Lingxiang Hu
Nanjing University
Jiacheng Huang
Jiacheng Huang
Nanjing University
entity resolutionknowledge graphcrowdsourcing
Z
Ziqian Jiang
Nanjing University
J
Jingkai Luo
Nanjing University
W
Wei Gao
Nanjing University
Y
Yunfan Tan
Hong Kong Polytechnic University
Z
Zun Wang
Tencent
J
Jiaheng Liu
Nanjing University