Less is More: DocString Compression in Code Generation

📅 2026-04-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the problem of excessive prompt length and high inference cost in code generation caused by redundant DocStrings, this paper proposes ShortenDoc, a task-specific DocString compression method. Unlike generic compression techniques—which achieve only ~10% compression—ShortenDoc is the first approach to jointly leverage code functionality intent recognition and structure-aware strategies. It employs an iterative pruning-and-rewriting mechanism constrained by semantic fidelity, achieving high compression rates of 25–40% without degrading code generation quality. Extensive experiments across six code generation benchmarks, five open-source LLMs (1B–10B parameters), and GPT-4o demonstrate that ShortenDoc significantly reduces token overhead while preserving code accuracy at baseline levels. This breaks the performance bottleneck of existing methods, establishing a new state-of-the-art for efficient DocString-aware code generation.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageData Mining & Knowledge Management: Data CompressionMachine Learning: Learning on the Edge & Model Compression

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
The widespread use of Large Language Models (LLMs) in software engineering has intensified the need for improved model and resource efficiency. In particular, for neural code generation, LLMs are used to translate function/method signature and DocString to executable code. DocStrings which capture user re quirements for the code and used as the prompt for LLMs, often contains redundant information. Recent advancements in prompt compression have shown promising results in Natural Language Processing (NLP), but their applicability to code generation remains uncertain. Our empirical study show that the state-of-the-art prompt compression methods achieve only about 10% reduction, as further reductions would cause significant performance degradation. In our study, we propose a novel compression method, ShortenDoc, dedicated to DocString compression for code generation. Our extensive experiments on six code generation datasets, five open-source LLMs (1B to 10B parameters), and one closed-source LLM GPT-4o confirm that ShortenDoc achieves 25-40% compression while preserving the quality of generated code, outperforming other baseline methods at similar compression levels. The benefit of this research is to improve efficiency and reduce the cost while maintaining the quality of the generated code, especially when calling third-party APIs, and is able to reduce the token processing cost by 25-40%.
Problem

Research questions and friction points this paper is trying to address.

Compress DocStrings for efficient code generation
Reduce token costs without degrading code quality
Optimize LLM prompts by removing redundant information
Innovation

Methods, ideas, or system contributions that make the work stand out.

DocString compression for code generation
25-40% compression with quality preservation
Reduces token processing cost significantly
🔎 Similar Papers
No similar papers found.
Nanjing University of Aeronautics and Astronautics | Nantong University | Monash University | National University of Defense Technology | Singapore Management University | Birkbeck, University of London
G
Guang Yang
Nanjing University of Aeronautics and Astronautics, China
Y
Yu Zhou
Nanjing University of Aeronautics and Astronautics, China
W
Wei Cheng
Nanjing University of Aeronautics and Astronautics, China
X
Xiangyu Zhang
Nanjing University of Aeronautics and Astronautics, China
X
Xiang Chen
Nantong University, China
Terry Yue Zhuo
Terry Yue Zhuo
Researcher
Large Language ModelsCode GenerationAI4SECybersecurity
K
Ke Liu
National University of Defense Technology, China
X
Xin Zhou
Singapore Management University, Singapore
D
David Lo
Singapore Management University, Singapore
Taolue Chen
Taolue Chen
School of Computing and Mathematical Sciences, Birkbeck, University of London
Software EngineeringProgram Analysis and VerificationMachine learning