Automated and Context-Aware Code Documentation Leveraging Advanced LLMs

📅 2025-09-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Prior automated code documentation generation research primarily targets code summarization, lacking context-aware approaches tailored to templated documentation (e.g., Javadoc) and suffering from the absence of high-quality, modern Java– and mainstream-framework–inclusive datasets. Method: We introduce the first context-aware Javadoc generation dataset, explicitly incorporating class/method signatures, call-site context, framework-specific APIs, and structured semantic information. Leveraging this dataset, we systematically evaluate open-source LLMs—including LLaMA-3.1, Gemma-2, Phi-3, Mistral, and Qwen-2.5—under zero-shot, few-shot, and fine-tuning paradigms. Contribution/Results: Experiments demonstrate that LLaMA-3.1 achieves consistently superior and robust performance across all settings, empirically validating the critical role of contextual modeling in templated documentation generation. Our dataset establishes a reproducible foundation for industrial-grade intelligent documentation systems, and our evaluation framework provides a practical technical pathway for future research and deployment.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageComputer Vision: Large Vision Models

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Code documentation is essential to improve software maintainability and comprehension. The tedious nature of manual code documentation has led to much research on automated documentation generation. Existing automated approaches primarily focused on code summarization, leaving a gap in template-based documentation generation (e.g., Javadoc), particularly with publicly available Large Language Models (LLMs). Furthermore, progress in this area has been hindered by the lack of a Javadoc-specific dataset that incorporates modern language features, provides broad framework/library coverage, and includes necessary contextual information. This study aims to address these gaps by developing a tailored dataset and assessing the capabilities of publicly available LLMs for context-aware, template-based Javadoc generation. In this work, we present a novel, context-aware dataset for Javadoc generation that includes critical structural and semantic information from modern Java codebases. We evaluate five open-source LLMs (including LLaMA-3.1, Gemma-2, Phi-3, Mistral, Qwen-2.5) using zero-shot, few-shot, and fine-tuned setups and provide a comparative analysis of their performance. Our results demonstrate that LLaMA 3.1 performs consistently well and is a reliable candidate for practical, automated Javadoc generation, offering a viable alternative to proprietary systems.
Problem

Research questions and friction points this paper is trying to address.

Addressing the gap in automated template-based documentation generation using LLMs
Overcoming the lack of Javadoc-specific datasets with modern language features
Evaluating public LLMs for context-aware Javadoc generation in Java codebases
Innovation

Methods, ideas, or system contributions that make the work stand out.

Context-aware dataset for Javadoc generation
Evaluated five open-source LLMs setups
LLaMA 3.1 reliable for automated documentation
🔎 Similar Papers
No similar papers found.
S
Swapnil Sharma Sarker
Ahsanullah University of Science and Technology, Dhaka, Bangladesh
T
Tanzina Taher Ifty
George Mason University, Fairfax, Virginia, USA