extract abstract syntax trees

Designs and implements tools that parse source code into abstract syntax trees and extract AST-based artifacts such as method and class nodes, documentation nodes, and associated metadata. Builds and analyzes these AST representations to compute static code metrics and signals and to detect code patterns or indicators relevant to vulnerabilities and other code-quality analyses.

extractabstractsyntaxtrees

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Code cloning is a major contributor to high maintenance costs and security risks; however, existing AST-based deep learning approaches suffer from insufficient semantic representation. Method: This paper systematically evaluates the effectiveness of various graph representations—including ASTs, CFGs, DFGs, and FA-ASTs—and their combinations, in conjunction with GNN architectures (GCN, GAT, GMN) for code clone detection. Contribution/Results: We reveal a strong coupling between graph fusion strategies and model architecture: AST+CFG+DFG significantly improves accuracy for GCN and GAT, whereas FA-AST degrades performance due to structural redundancy; remarkably, GMN achieves superior performance using AST alone, outperforming most fused variants. Our approach attains 98.2% accuracy across multiple benchmarks. This work establishes, for the first time, the optimal matching relationships between GNN architectures and graph representations, providing a reusable, principled modeling guideline for industrial-scale code clone detection.

Assessing compatibility of enriched ASTs with GNN architecturesDetermining optimal graph structures for accurate clone detectionEvaluating hybrid AST-graph representations for code clone detection

Existing algorithm identification methods often suffer from poor usability, limited scalability, and insufficient evaluation. This work proposes a novel paradigm that integrates domain-specific languages (DSLs) with abstract syntax tree (AST) pattern matching: algorithmic characteristics are formally specified using a DSL to construct a reusable library of AST patterns, enabling automatic recognition of common algorithm implementations in source code. Evaluated on a subset of BigCloneEval, the approach achieves an average F1 score of 0.74, substantially outperforming CodeLlama (0.35) and state-of-the-art code clone detectors, which attain a recall of only 0.20 compared to our method’s 0.62. This advance represents a dual improvement in both precision and practical applicability for algorithm identification.

abstract syntax treealgorithm recognitionautomated analysis

Current large language models (LLMs) lack software engineering context when generating code explanations, leading to hallucinations and limiting their practical utility in code maintenance, developer onboarding, and legacy system modernization. To address this, we propose the first systematic framework that leverages multi-source natural language artifacts from GitHub—including pull request descriptions, issue discussions, and commit messages—to enhance code intent understanding. Our approach comprises three tightly integrated components: contextual artifact extraction, high-level explanation generation, and automated verification. By employing structured context modeling and integrating the Model Context Protocol (MCP), we enable LLM-driven code explanations that are both high-fidelity and verifiable. Empirical evaluation demonstrates a substantial improvement in explanation accuracy and a near-zero hallucination rate. Both open-source contributors and enterprise developers confirm that the generated insights deliver tangible value in collaborative software development practices.

Enhancing code understanding by leveraging GitHub artifacts for contextGenerating grounded code explanations using pull requests and issuesReducing LLM hallucinations in code analysis through structured validation

AI-Driven Code Refactoring: Using Graph Neural Networks to Enhance Software Maintainability

Apr 14, 2025
GB
Gopichand Bandarupalli
🏛️ Campbellsville University

This work addresses the declining maintainability of software caused by high cyclomatic complexity and coupling in code refactoring. We propose the first end-to-end Graph Neural Network (GNN)-driven, semantics-aware refactoring method. By modeling Abstract Syntax Trees (ASTs) as graphs and integrating AST embeddings with static analysis, our approach automatically identifies and optimizes high-complexity, high-coupling code fragments. Unlike conventional rule-based or shallow-model approaches, ours is the first to systematically apply GNNs across the entire refactoring decision pipeline. Evaluated on 2 million Python code snippets, our method achieves 92% refactoring accuracy, reduces average cyclomatic complexity by 35%, and decreases coupling by 33%. These improvements significantly outperform established baselines—including SonarQube and decision tree–based methods—demonstrating both technical novelty and practical efficacy in automated, semantics-guided code refactoring.

Comparing GNN performance against rule-based and decision tree methodsReducing code complexity and coupling with AI-driven refactoringUsing GNNs to improve software maintainability via code refactoring

Code Arcades: 3d Visualization of Classes, Dependencies and Software Metrics

Sep 27, 2025
AS
Anthony Savidis
🏛️ University of Crete | Alpha Omega Zed SA

This study addresses the insufficient structural information visualization in understanding and maintaining complex software systems. We propose a configurable three-dimensional (3D) visualization method that integrates fine-grained metrics (e.g., cyclomatic complexity, coupling) with coarse-grained ones (e.g., module- or package-level abstractions) to construct an interactive 3D rendering engine supporting dynamic attribute adjustment. Our approach introduces a novel configurable code-element grouping mechanism and synergistically incorporates dependency structure graphs, treemaps, and version history to enable multi-scale visualization—from macro-level metaphorical overviews to micro-level focused analysis. Empirical evaluation demonstrates that our method significantly improves efficiency in identifying structural patterns and localizing architectural hotspots. Compared to conventional two-dimensional visualization techniques, it enhances both the depth of system comprehension and maintenance effectiveness—particularly in large-scale codebases.

Identifying complexity hotspots and system behavior patternsProviding multi-level metrics and interactive code explorationVisualizing software structure and dependencies graphically

Latest Papers

What's happening recently
View more

MLCPD: A Unified Multi-Language Code Parsing Dataset with Universal AST Schema

Oct 18, 2025
JG
Jugal Gajjar
🏛️ The George Washington University

Existing code datasets are predominantly confined to single-language lexical features or isolated parsers, hindering cross-lingual syntactic reasoning and structural analysis. To address this, we propose a language-agnostic, universal Abstract Syntax Tree (AST) abstraction schema that enables structural alignment and semantic normalization across ten mainstream programming languages. We introduce the first large-scale, high-fidelity multilingual code parsing dataset—comprising over 7 million source files—generated via a unified compilation pipeline and stored in Parquet format, accompanied by reproducibility scripts and interactive visualization tools. The dataset is publicly released on Hugging Face and GitHub. Empirical analysis reveals substantial syntactic structural commonalities across languages, providing foundational support for cross-lingual program understanding, pretraining of code models, and static program analysis.

Enabling consistent cross-language reasoning and structural learningProviding hierarchical tree structures with universal AST schemaUnifying syntactic code representations across ten programming languages

ATLAS: Automated Tree-based Language Analysis System for C and C++ source programs

Dec 13, 2025
JM
Jaid Monwar Chowdhury
🏛️ Bangladesh University of Engineering and Technology

Static analysis of C/C++ programs faces significant challenges due to pointer aliasing, multi-level indirection, function pointers, and type ambiguity induced by `typedef`. To address these, this paper proposes an end-to-end, compiler-agnostic, interprocedural, type-aware static analysis system. Methodologically, it leverages Clang LibTooling to construct a unified multi-view intermediate representation—integrating AST, CFG, and DFG—and introduces custom CFG/DFG construction algorithms alongside an alias-aware type inference module, enabling the first interprocedural, type-sensitive data-flow modeling. Contributions include: (1) statement-level control-flow and type-aware data-flow graph generation for uncompiled code; (2) empirical validation on real-world open-source projects demonstrating high-fidelity modeling of complex semantic dependencies; and (3) provision of interpretable, structurally enriched input representations for downstream software engineering tasks such as vulnerability detection and code completion.

Generates control and data flow graphs for C/C++ programsHandles compilable and non-compilable multi-file C/C++ projectsProduces unified multi-view code representations for program analysis

This work proposes a unified framework that integrates graph neural networks with large language models (LLMs) to jointly detect, explain, and repair software maintainability and security issues. Addressing the high false-positive rates and maintenance overhead of existing code smell and vulnerability detection tools—stemming from their lack of structured contextual awareness—the approach uniquely fuses multi-dimensional program graphs, including abstract syntax trees (ASTs), control flow graphs (CFGs), and program dependence graphs (PDGs), with deep code embeddings. The resulting model is cross-lingual, interpretable, and readily integrable into CI/CD pipelines. Empirical evaluation on multilingual datasets demonstrates significant improvements over conventional rule-based analyzers and single-model baselines, achieving higher detection accuracy and generating more practical repair suggestions.

AI-assisted code reviewcode smellsprogram analysis

This study addresses the challenge of automatically detecting software design patterns in source code to support architectural understanding and quality assessment. It presents the first systematic evaluation of four large language models—including NextCoder and Gemma 3—as well as two ensemble strategies combining three models, for recognizing five classic design patterns: Singleton, Adapter, Bridge, Composite, and Decorator. The work investigates the impact of three input modalities—raw source code, PlantUML diagrams, and textual descriptions—on detection performance. Experimental results demonstrate that NextCoder and Gemma 3 achieve the highest accuracy among individual models, while ensemble approaches further enhance performance, thereby confirming the effectiveness and potential of large language models in design pattern recognition tasks.

automatic detectiondesign pattern recognitionsoftware architecture understanding

This work proposes AST(NIT), a novel approach that leverages fully serialized abstract syntax trees (ASTs) to enhance code summarization with large language models (LLMs). While existing methods typically rely on raw source code or partial AST information, AST(NIT) encodes complete structural details of the AST while preserving lexical content, producing compact input sequences tailored for LLMs. In the first systematic evaluation of its kind, the method is integrated with the LLaMA-3.1-8B model on the CodeXGLUE Python dataset, demonstrating significantly reduced input length and training time without compromising summary quality. The results confirm that fully serialized ASTs offer both effectiveness and efficiency advantages in LLM-based code summarization tasks.

Abstract Syntax TreeAST serializationcode summarization

Hot Scholars

YZ

Yanjie Zhao

Huazhong University of Science and Technology
Software EngineeringSoftware Security
ML

Mingwei Liu

Rutgers University
China laborhigh performance work systems
AH

Andre Hora

Universidade Federal de Minas Gerais (UFMG)
Software EvolutionSoftware MaintenanceSoftware TestingMining Software Repositories
HL

Huawei Li

Institute of Computing Technology, Chinese Academy of Sciences
computer engineering