Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation: An Empirical Study

📅 2024-11-04
🏛️ arXiv.org
📈 Citations: 5
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitively high computational cost of full-parameter fine-tuning for large language models (LLMs) in unit test generation—a key barrier to practical deployment. We conduct the first systematic empirical evaluation of three parameter-efficient fine-tuning (PEFT) methods—LoRA, (IA)³, and Prompt Tuning—across multiple LLM scales. On standard code testing benchmarks, all three approaches achieve generation quality comparable to full fine-tuning while updating fewer than 0.5% of model parameters. Prompt Tuning delivers the best trade-off between resource efficiency and effectiveness, yielding the lowest overall cost; LoRA most closely matches full fine-tuning performance across diverse scenarios; and (IA)³ exhibits robustness but slower convergence. Our work fills a critical empirical gap in applying PEFT to automated software testing and provides a reproducible methodology and practical guidance for lightweight, LLM-driven test generation.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageSearch and Optimization: Learning to Search

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Web evaluation methodologies and metrics
📝 Abstract
The advent of large language models (LLMs) like GitHub Copilot has significantly enhanced programmers' productivity, particularly in code generation. However, these models often struggle with real-world tasks without fine-tuning. As LLMs grow larger and more performant, fine-tuning for specialized tasks becomes increasingly expensive. Parameter-efficient fine-tuning (PEFT) methods, which fine-tune only a subset of model parameters, offer a promising solution by reducing the computational costs of tuning LLMs while maintaining their performance. Existing studies have explored using PEFT and LLMs for various code-related tasks and found that the effectiveness of PEFT techniques is task-dependent. The application of PEFT techniques in unit test generation remains underexplored. The state-of-the-art is limited to using LLMs with full fine-tuning to generate unit tests. This paper investigates both full fine-tuning and various PEFT methods, including LoRA, (IA)^3, and prompt tuning, across different model architectures and sizes. We use well-established benchmark datasets to evaluate their effectiveness in unit test generation. Our findings show that PEFT methods can deliver performance comparable to full fine-tuning for unit test generation, making specialized fine-tuning more accessible and cost-effective. Notably, prompt tuning is the most effective in terms of cost and resource utilization, while LoRA approaches the effectiveness of full fine-tuning in several cases.
Problem

Research questions and friction points this paper is trying to address.

Investigating parameter-efficient fine-tuning methods for unit test generation using large language models
Evaluating performance trade-offs between computational cost and test generation effectiveness
Analyzing syntax correctness and coverage metrics of generated unit tests
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parameter-efficient fine-tuning reduces computational costs
LoRA achieves comparable performance to full fine-tuning
Prompt tuning is most cost-effective for large models