PreMaQ: Predicting Maintainability-Related Quality of LLM-Generated Code Before Generation

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that large language models (LLMs) struggle to estimate code maintainability prior to generation, often producing low-quality code that increases review costs. To this end, this work proposes the first pre-generation prediction method for code maintainability based on LLM internal representations. By leveraging prompt embedding techniques for regression, it estimates smell scores and maintainability indices before code generation to optimize model selection. The authors demonstrate that this maintainability prediction is complementary to functional correctness prediction. Experimental results show Spearman correlation coefficients of 0.57 and 0.65, with model selection efficacy reaching 60.7% of the ideal performance.
📝 Abstract
As large language models (LLMs) become increasingly capable of code generation, adopting generated code in software development requires assessing not only its functional correctness but also its maintainability-related quality. If such quality could be estimated before generation, developers could avoid the cost of generating, reviewing, and discarding low-quality code. Although prior work has shown that the functional correctness of the LLM-generated code can be predicted in advance, it remains unclear whether maintainability-related quality is similarly predictable. We introduce Pre-Generation Maintainability-Related Quality Prediction (PreMaQ), which predicts the Code Smell Score (CSS) and Maintainability Index (MI) of generated code from the internal representations of LLMs before generation. Our evaluation covers four open-weight LLMs and four Python code generation benchmarks, comprising 2,695 tasks in total. Our results show that predicted CSS and MI consistently correlate with their observed values across all 16 model-benchmark combinations, achieving mean Spearman rank correlations of 0.57 and 0.65, respectively. When used for model selection, PreMaQ achieves 59.5% of the maintainability-related quality improvement attainable by an ideal maintainability-based selector over random selection on tasks for which multiple models generate functionally correct code. Combining predictions from PreMaQ and prompt embeddings increases this proportion to 60.7%, indicating that the two signals are complementary. These findings suggest that PreMaQ can be used to predict and improve the maintainability-related quality of LLM-generated code.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Code Generation
Maintainability
Quality Prediction
Code Smell
Innovation

Methods, ideas, or system contributions that make the work stand out.

Maintainability Prediction
Pre-Generation
Internal Representations
Model Selection
Code Quality
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.