Reinforcement Learning from Automatic Feedback for High-Quality Unit Test Generation

📅 2026-04-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Unit tests generated by large language models (LLMs) frequently contain errors and code smells, primarily due to contamination from low-quality test patterns in their training data. Method: We propose RLSQM, a reinforcement learning framework that leverages static quality metrics to guide LLM fine-tuning. RLSQM is the first approach to formulate fine-grained static indicators—such as maintainability and correctness—as differentiable, trainable reward signals. It employs Proximal Policy Optimization (PPO) for stepwise optimization and integrates these signals into a unified multi-objective reward model, combining static code analysis, reward modeling, and LLM adaptation. Results: Experiments show RLSQM improves test quality by up to 21% over baseline LLMs, achieves near-perfect syntactic correctness (~100%), and outperforms GPT-4 on 4 of 7 quantitative metrics. The method significantly enhances the reliability and engineering readiness of generated unit tests.
📝 Abstract
Software testing is a crucial aspect of software development, and the creation of high-quality tests that adhere to best practices is essential for effective maintenance. Recently, Large Language Models (LLMs) have gained popularity for code generation, including the automated creation of test cases. However, these LLMs are often trained on vast amounts of publicly available code, which may include test cases that do not adhere to best practices and may even contain test smells (anti-patterns). To address this issue, we propose a novel technique called Reinforcement Learning from Static Quality Metrics (RLSQM). To begin, we analyze the anti-patterns generated by the LLM and show that LLMs can generate undesirable test smells. Thus, we train specific reward models for each static quality metric, then utilize Proximal Policy Optimization (PPO) to train models for optimizing a single quality metric at a time. Furthermore, we amalgamate these rewards into a unified reward model aimed at capturing different best practices and quality aspects of tests. By comparing RL-trained models with those trained using supervised learning, we provide insights into how reliably utilize RL to improve test generation quality and into the effects of various training strategies. Our experimental results demonstrate that the RL-optimized model consistently generated high-quality test cases compared to the base LLM, improving the model by up to 21%, and successfully generates nearly 100% syntactically correct code. RLSQM also outperformed GPT-4 on four out of seven metrics. This represents a significant step towards enhancing the overall efficiency and reliability of software testing through Reinforcement Learning and static quality metrics. Our data are available at https://figshare.com/s/ded476c8d4c221222849.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Unit Test Case Generation
Quality Assurance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Large Language Models
Software Testing Optimization
🔎 Similar Papers
No similar papers found.