Less is More: Towards Green Code Large Language Models via Unified Structural Pruning

📅 2026-04-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Generative large language models (LLMs) for code suffer from high energy consumption and computational overhead. Existing structural pruning methods—designed primarily for classification tasks—exhibit objective mismatch and component fragmentation when applied to high-dimensional token sequence generation. Method: This paper proposes Flab-Pruner, a unified structural pruning framework featuring a novel “vocabulary–layer–FFN” ternary co-pruning paradigm. It integrates parameter sensitivity analysis with code-domain-specific instruction fine-tuning to overcome classification-oriented pruning limitations. Contribution/Results: Applied to three state-of-the-art code LLMs, Flab-Pruner achieves 22% parameter compression while retaining 97% of original performance; post-training yields comparable or superior accuracy. It significantly reduces GPU memory footprint, power consumption, and carbon emissions, demonstrating strong deployment robustness and practical efficiency for sustainable code generation.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageMachine Learning: Large Multimodal Models (LMMs)Computer Vision: Large Vision Models

Application Category

Search and Retrieval-Augmented AI: Large language models for searchEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
The extensive application of Large Language Models (LLMs) in generative coding tasks has raised concerns due to their high computational demands and energy consumption. Unlike previous structural pruning methods designed for classification models that deal with lowdimensional classification logits, generative Code LLMs produce high-dimensional token logit sequences, making traditional pruning objectives inherently limited. Moreover, existing single component pruning approaches further constrain the effectiveness when applied to generative Code LLMs. In response, we propose Flab-Pruner, an innovative unified structural pruning method that combines vocabulary, layer, and Feed-Forward Network (FFN) pruning. This approach effectively reduces model parameters while maintaining performance. Additionally, we introduce a customized code instruction data strategy for coding tasks to enhance the performance recovery efficiency of the pruned model. Through extensive evaluations on three state-of-the-art Code LLMs across multiple generative coding tasks, the results demonstrate that Flab-Pruner retains 97% of the original performance after pruning 22% of the parameters and achieves the same or even better performance after post-training. The pruned models exhibit significant improvements in storage, GPU usage, computational efficiency, and environmental impact, while maintaining well robustness. Our research provides a sustainable solution for green software engineering and promotes the efficient deployment of LLMs in real-world generative coding intelligence applications.
Problem

Research questions and friction points this paper is trying to address.

Reducing computational demands of Code LLMs via structural pruning
Overcoming limitations of traditional pruning for generative tasks
Maintaining model performance while improving efficiency and sustainability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified structural pruning for Code LLMs
Combines vocabulary, layer, and FFN pruning
Customized code instruction data strategy
🔎 Similar Papers
No similar papers found.