Code Health in LLM-Based Test Generation: Effectiveness and Token Efficiency

📅 2026-08-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了代码可维护性对LLM生成单元测试有效性的影响,使用CodeHealth指标评估,并发现其与输入token数负相关。
📝 Abstract
Coding agents powered by Large Language Models (LLMs) are now prominent in software engineering. Previous work has shown that AI tools perform better on high-quality source code that is easy to maintain. In this study, we investigate how the effectiveness of LLM-generated unit tests varies across maintainability levels measured by CodeScene's CodeHealth (CH). We assess test effectiveness using traditional coverage metrics and mutation score across Python, Java, and C++. Moreover, we study how code with different levels of CH translates into input tokens using common industrial tokenizers. Our results suggest that CH provides a weak but consistent signal of LLM-generated test effectiveness and is negatively correlated with input-token count. These findings provide further evidence for a relationship between maintainability and LLM-based software development.
Problem

Research questions and friction points this paper is trying to address.

Code Health
LLM-based Test Generation
Maintainability
Token Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
CodeHealth
test generation
token efficiency
maintainability
F
Freya Wirdemann
Heidelberg University
M
Markus Borg
CodeScene and Lund University
N
Nadim Hagatulah
Lund University
A
Adam Tornhill
CodeScene