parse code diffs

Designs and implements parsers and processing pipelines that convert textual code diffs into structured representations (files, hunks, and lines), map tokens back to their original code lines, and aggregate token-level analyses into higher-level units while preserving the mappings relied on by a developer’s mental model.

parsecodediffs

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing direct code-to-code transformation approaches often suffer from semantic drift, implicit behavioral changes, and loss of traceability. To address these issues, this work proposes a specification-based Code2Text2Code refactoring framework that first translates source code into a neutral textual specification before generating target code. The approach integrates abstract syntax tree (AST) and dependency graph analysis, semantic-aware code chunking, retrieval-augmented generation, and DSPy-based prompt tuning, further enhanced by iterative validation and graph-based formal verification. This pipeline ensures high-fidelity semantic preservation and controllable evolution during code transformation. Experimental results demonstrate that the proposed method significantly reduces transformation loss and substantially improves semantic consistency, interface stability, and cross-language traceability of the refactored code.

behavioral changesCode2Code transformationdomain logic reconstruction

On the Effect of Token Merging on Pre-trained Models for Code

Jul 18, 2025
MS
Mootez Saad
🏛️ Dalhouise University | Queen’s University

This work addresses the computational overhead induced by subword representation redundancy in code pre-trained models. We propose two semantic-driven hidden-layer representation merging strategies: static averaging and dynamic fusion via learnable weights. To our knowledge, this is the first systematic investigation into merging hidden representations of subwords belonging to the same semantic unit—e.g., constituent subwords of a single identifier—while preserving identifier-level semantic integrity. Our approach integrates seamlessly into mainstream models including CodeBERT, UniXCoder, and CodeT5+, and is empirically validated on vulnerability detection, code classification, and code translation tasks. Experiments demonstrate a 1–19% reduction in inference computation, a +2.47 improvement in CodeBLEU for code translation, and only a marginal −1.82 drop in F1-score for vulnerability detection—achieving a Pareto improvement in both efficiency and performance.

Evaluates methods on multiple code tasks and modelsInvestigates token merging impact on code language modelsProposes strategies to reduce computational overhead in tokenization

Large language models (LLMs) excel at code generation, yet the impact of their compressed variants—such as those produced via quantization or knowledge distillation—on token representations of programming languages remains poorly understood, hindering deployment quality. This work systematically investigates how LLM tokenizers encode programming languages and introduces a novel “cold-start probability” analysis method that operates without explicit prompting. By integrating lexical distribution analysis, keyword coverage, and multidimensional evaluation metrics, the study provides the first comprehensive characterization of the subtle effects of compression strategies—including quantization, knowledge distillation, model scaling, and task-specific fine-tuning—on code token representations. The findings offer both theoretical grounding and empirical guidance for deploying high-quality, efficient code generation models.

code generationcompressed LLMsdistillation

This study presents the first systematic evaluation of Transformer models’ robustness under semantics-preserving code transformations. To address this, we construct a benchmark for Java and Python comprising 51 transformation strategies—categorized into five perturbation types—and evaluate performance across three fundamental code intelligence tasks: code completion, code summarization, and code retrieval. Methodologically, our approach integrates abstract syntax tree (AST)-based structural representations, multi-granularity transformations (including block-level edits, insertions/deletions, syntactic/token-level modifications, and identifier replacements), and comparative analysis of positional encoding schemes. Key findings include: (1) AST-aware encoding substantially improves robustness, yielding absolute accuracy gains of 12–28%; (2) insertion/deletion and identifier-level transformations are the most adversarial; and (3) positional encoding design critically modulates model resilience. This work establishes an empirically grounded, reproducible benchmark and provides actionable insights for robustness-aware modeling and architectural refinement of code intelligence systems.

Analyze impact of semantic-preserving changes on code intelligence tasksCompare AST-based vs sequence-based Transformer performance on perturbed codeStudy robustness of Transformer models under code transformations

Latest Papers

What's happening recently
View more

This work addresses the challenge of identifying structural and semantic similarities across imperative programs written in different languages by proposing a unified graph representation that integrates abstract syntax trees with neural semantic embeddings. The approach transforms annotated programs into typed, attributed graphs and leverages CodeBERT and SentenceTransformer to generate rich semantic embeddings. By constructing consistent graph representations across multilingual verification datasets—including C/ACSL, Java/JML, and Dafny—it achieves, for the first time, joint modeling of syntactic structure and formal semantics. This unified framework offers a viable pathway for cross-language reuse of verification artifacts and demonstrates strong generality and effectiveness across diverse programming languages and specification frameworks.

graph constructionimperative programsprogram representation

This study addresses the frequent failure of large language models (LLMs) in software engineering tasks due to structurally invalid outputs—such as syntactic or formatting errors—that prevent correct parsing by downstream toolchains, even when the semantic content is accurate. The authors systematically evaluate the reliability of structured output generation across four representative tasks, categorizing errors into syntactic, structural, and semantic types. They propose a template-driven token-matching generation (TTMG) method that enforces structural consistency during autoregressive decoding. Experimental results demonstrate that TTMG nearly eliminates syntactic errors; however, structural and semantic errors remain prevalent. This work reveals, for the first time, that the fundamental bottleneck in structured output generation lies not merely in syntax but in the insufficient coordination between structural and semantic correctness, indicating that existing structural control mechanisms, while necessary, are insufficient without joint guarantees of both dimensions.

format compliancelarge language modelssoftware engineering

Current large code models exhibit limited performance in repository-level code generation due to their neglect of cross-file dependencies and structural context, while conventional NLP-based retrieval-augmented approaches struggle to effectively model the inherent structure of code. To address this, this work proposes Hydra, a novel framework that treats code as structured entities and introduces a hierarchical code tree index, a dependency-aware retriever (DAR), and a hybrid retrieval mechanism that integrates functional dependencies with semantic similarity. Hydra departs from traditional NLP-style paradigms for code processing and achieves state-of-the-art results on the DevEval and RepoExec benchmarks, surpassing the strongest baseline by over 5% in Pass@1. Notably, it enables smaller models equipped with Hydra to match the performance of larger models using conventional retrievers.

code coherencecode structurecross-file dependencies

This work addresses the prevalent issue in large language models (LLMs) of introducing control-flow, type, or I/O errors during code translation due to neglect of program intent. To mitigate this, the paper proposes the first systematic use of a language-agnostic, structured intermediate specification that preserves semantic fidelity through an intermediate representation, structured generation, and automated test-based validation. Evaluated on the Avatar and CodeNet datasets with five state-of-the-art LLMs, the approach significantly improves translation accuracy, raising the micro-averaged accuracy from 67.7% to 78.5%. It completely eliminates lexical errors and substantially reduces errors related to structure, declarations, and runtime dependencies.

code translationcross-language programmingLarge Language Models

This work addresses the challenge posed by frequent and large-scale code changes in modern software projects, which overwhelm traditional code review practices. While existing large language model (LLM)-based approaches primarily focus on generating summaries, they lack structured identification of change types. To bridge this gap, the paper proposes a two-stage pipeline that leverages LLMs to perform taxonomy-based structured labeling of code diffs and extract semantic relationships and attributes—such as rename propagation and type modifications. This approach represents the first systematic exploration of LLMs for structured understanding of code changes, operating without reliance on static analysis toolchains and supporting language-agnostic, customizable taxonomies. Evaluated on both natural and synthetic patch benchmarks, the best configuration achieves 84% recall and 81% precision, with notably high accuracy in metadata extraction.

code change labelingcode reviewlarge language models

Hot Scholars

CT

Christoph Treude

Associate Professor of Computer Science, Singapore Management University
Software EngineeringEmpirical Software EngineeringHuman-AI InteractionAI for Science
AM

Audris Mockus

University of Tennessee
Digital ArchaeologySoftware EngineeringVisualizationOptimization
PN

Pengyu Nie

University of Waterloo
Software EngineeringNatural Language ProcessingProgramming Languages
BL

Bach Le

Senior Lecturer, ARC DECRA, School of Computing and Information Systems, The University of Melbourne
Program RepairProgram SynthesisFormal MethodsComputer Security
HT

Haoye Tian

Assistant Professor, Aalto University
Software EngineeringMachine LearningProgram RepairAI4SE