Score
Designs and implements tools and pipelines that extract and assemble contextual information for software code elements (files, classes, and methods), including project structure and metadata, build and classpath resolution, source and documentation text, and static call‑graph relationships; produces method‑level context records that combine these diverse artifacts for downstream analysis and tooling.
Code intelligence research has long lacked a systematic analysis of context utilization; although prior work confirms context improves model performance, context types remain ambiguously defined, integration strategies are fragmented, and evaluation protocols are inconsistent. Method: We conduct a comprehensive survey of 146 studies published between 2007 and 2024, introducing the first context taxonomy specifically designed for code intelligence, and establishing a three-dimensional analytical framework—“task–context–evaluation.” Using bibliometric analysis, systematic literature review, and qualitative coding, we rigorously synthesize empirical evidence. Contributions: (1) A quantitative developmental map of the field; (2) A novel, principled context classification system; (3) A comparative analysis of context integration strategies across 12 code intelligence tasks; and (4) A diagnostic assessment of prevalent evaluation shortcomings, accompanied by a practical, actionable research roadmap.
This work addresses the high cost, inconsistency, and poor reproducibility associated with manual collection of method-level contextual information—such as class metadata, documentation, and call relationships—in large-scale Java projects. To overcome these challenges, the authors propose the first task-agnostic, reusable unified pipeline that automatically parses Maven/Gradle project structures and classpaths, leverages SootUp to construct static call graphs, employs Spoon for source code analysis, and achieves precise alignment between source code and bytecode to generate a versioned, multidimensional context dataset. Evaluated on 20 real-world repositories, the pipeline successfully processes 56,512 methods and 386,048 call edges, with 97.8% of intra-project call edges accurately mapped to source code locations and a human-audited correctness rate of 99.0%.
This work addresses the limited cross-file contextual awareness of code large language models (CodeLLMs) in repository-level code generation. Methodologically, we introduce RepoExec—the first executable and functionally correct repository-level benchmark—and propose Dependency Invocation Rate (DIR), a novel metric quantifying the accuracy of cross-file dependency invocation. We further design an instruction-tuning dataset integrating test-driven validation and context-aware dependency modeling. Our contributions include the first comprehensive evaluation framework encompassing context-awareness, execution-driven assessment, and cross-file dependency modeling. Experimental results demonstrate that instruction tuning significantly improves contextual utilization and debugging capability, whereas pre-trained models exhibit stronger functional correctness. RepoExec has since become the de facto standard benchmark for repository-level code generation research.
Prior automated code documentation generation research primarily targets code summarization, lacking context-aware approaches tailored to templated documentation (e.g., Javadoc) and suffering from the absence of high-quality, modern Java– and mainstream-framework–inclusive datasets. Method: We introduce the first context-aware Javadoc generation dataset, explicitly incorporating class/method signatures, call-site context, framework-specific APIs, and structured semantic information. Leveraging this dataset, we systematically evaluate open-source LLMs—including LLaMA-3.1, Gemma-2, Phi-3, Mistral, and Qwen-2.5—under zero-shot, few-shot, and fine-tuning paradigms. Contribution/Results: Experiments demonstrate that LLaMA-3.1 achieves consistently superior and robust performance across all settings, empirically validating the critical role of contextual modeling in templated documentation generation. Our dataset establishes a reproducible foundation for industrial-grade intelligent documentation systems, and our evaluation framework provides a practical technical pathway for future research and deployment.
Large language models (LLMs) face two key challenges in code completion: limited context window capacity and susceptibility to noisy, irrelevant context. This paper investigates how context granularity—file-level versus block-level—and retrieval ranking strategies affect generation quality. We propose a static-analysis-driven, block-level context retrieval method that enables fine-grained, semantically relevant context extraction, followed by optimized context composition and ordering. Experiments on Python code completion show that our approach improves completion accuracy by 6% over the best-performing file-level retrieval baseline and by 16% over a no-context baseline. Our core contributions are: (1) empirical validation that block-level context significantly enhances effectiveness under strict context-length constraints; and (2) the first lightweight, deployable retrieval framework that jointly integrates static program analysis with context sequence control. This work establishes a practical, production-ready paradigm for context optimization in industrial code completion systems.
This work addresses the challenge of generating comprehensive UML diagrams from large-scale codebases, which is hindered by the context-length limitations of large language models (LLMs). To overcome this, the authors propose a hierarchical multi-agent architecture combined with a deterministic importance-weighted intermediate representation (IR) compression method that reduces massive codebases into context-appropriate views within milliseconds—without requiring LLM inference. Built upon the Claude Agent SDK, five specialized agents collaboratively enable cross-language, scalable, and automated UML generation. Evaluation across 12 open-source projects spanning four programming languages and seven UML diagram types demonstrates strong performance: a syntactic correctness rate of 91.5%, a mean relationship precision of 0.858, and an average structural quality score of 81.7 out of 100, with no degradation in performance as codebase size increases.
This work addresses the inaccuracy and limited utility of developer documentation generated by large language models (LLMs), which often stems from overlooking cross-file dependency chains. To mitigate this, the authors propose Context-as-a-Service (CaaS), a retrieval layer for LLM agents that integrates hybrid keyword and semantic search into the documentation generation pipeline. CaaS explicitly supports the discovery of non-obvious cross-file dependencies by jointly indexing source code, API references, and upstream documentation. Integrated with Claude Sonnet 4.6, CaaS identified four documentation issues and four tutorial flaws missed by baseline methods across two case studies, reduced task completion time by 22%–34%, and substantially decreased input token consumption.
This work addresses the limitation of existing code generation models that ignore calling context, often producing functions incompatible with real-world projects. The authors propose CallerGen, the first calling-aware code pre-training framework, which leverages static analysis to extract caller-callee pairs from real codebases and formulates a novel pre-training objective conditioned on calling context. They also introduce CallerEval, a dedicated benchmark for evaluating context-aware code generation. By explicitly incorporating calling context throughout both pre-training and evaluation, CallerGen significantly outperforms same-sized baselines on CallerEval, achieving Pass@1 accuracy of 16.58% and 22.81% for its 220M and 0.5B parameter variants, respectively—performance comparable to substantially larger models.
This study addresses the longstanding fragmentation in software artifact traceability research, characterized by incomplete linkages, ambiguous techniques, and disconnected application contexts. Through a systematic literature review, it constructs the first comprehensive traceability landscape encompassing 22 artifact types and 23 relationship kinds, and introduces a technology decision map, a standardized evaluation benchmark, and a role-oriented dynamic path alignment framework. The work uncovers critical challenges: a pervasive code-centric bias, a reproducibility crisis stemming from only 37% of studies releasing open-source artifacts, and a significant adoption gap with 95% of proposed tools never deployed in industry. In response, it offers targeted strategies to bridge these gaps, establishing a unified knowledge foundation for future research and practical implementation in traceability.
This work addresses the problem of “context decay” in AI coding assistants, where configuration files such as CLAUDE.md or AGENTS.md become outdated due to code evolution, leading to behavioral drift in AI responses. To mitigate this issue, the study proposes adapting existing software documentation consistency detection techniques—originally designed for READMEs and wikis—to identify inconsistencies between AI configuration artifacts and their corresponding codebases. By repurposing established toolchains that verify alignment between documentation and source code, the approach enables automated detection of stale references in configuration files. An empirical evaluation across 356 open-source repositories reveals that 23.0% of projects contain outdated references, demonstrating the viability of leveraging traditional consistency-checking tools to detect context decay. This finding offers a practical technical pathway for maintaining reliable AI-assisted development environments through improved configuration hygiene.