A Closer Look into Transformer-Based Code Intelligence Through Code Transformation: Challenges and Opportunities

📅 2022-07-09
🏛️ arXiv.org
📈 Citations: 5
✨ Influential: 1
📄 PDF
🤖 AI Summary
This study presents the first systematic evaluation of Transformer models’ robustness under semantics-preserving code transformations. To address this, we construct a benchmark for Java and Python comprising 51 transformation strategies—categorized into five perturbation types—and evaluate performance across three fundamental code intelligence tasks: code completion, code summarization, and code retrieval. Methodologically, our approach integrates abstract syntax tree (AST)-based structural representations, multi-granularity transformations (including block-level edits, insertions/deletions, syntactic/token-level modifications, and identifier replacements), and comparative analysis of positional encoding schemes. Key findings include: (1) AST-aware encoding substantially improves robustness, yielding absolute accuracy gains of 12–28%; (2) insertion/deletion and identifier-level transformations are the most adversarial; and (3) positional encoding design critically modulates model resilience. This work establishes an empirically grounded, reproducible benchmark and provides actionable insights for robustness-aware modeling and architectural refinement of code intelligence systems.
📝 Abstract
Transformer-based models have demonstrated state-of-the-art performance in many intelligent coding tasks such as code comment generation and code completion. Previous studies show that deep learning models are sensitive to the input variations, but few studies have systematically studied the robustness of Transformer under perturbed input code. In this work, we empirically study the effect of semantic-preserving code transformation on the performance of Transformer. Specifically, 24 and 27 code transformation strategies are implemented for two popular programming languages, Java and Python, respectively. For facilitating analysis, the strategies are grouped into five categories: block transformation, insertion/deletion transformation, grammatical statement transformation, grammatical token transformation, and identifier transformation. Experiments on three popular code intelligence tasks, including code completion, code summarization and code search, demonstrate insertion/deletion transformation and identifier transformation show the greatest impact on the performance of Transformer. Our results also suggest that Transformer based on abstract syntax trees (ASTs) shows more robust performance than the model based on only code sequence under most code transformations. Besides, the design of positional encoding can impact the robustness of Transformer under code transformation. Based on our findings, we distill some insights about the challenges and opportunities for Transformer-based code intelligence.
Problem

Research questions and friction points this paper is trying to address.

Study robustness of Transformer models under code transformations
Analyze impact of semantic-preserving changes on code intelligence tasks
Compare AST-based vs sequence-based Transformer performance on perturbed code
Innovation

Methods, ideas, or system contributions that make the work stand out.

Empirical study on Transformer robustness via code transformations
24 Java and 27 Python semantic-preserving transformation strategies
AST-based Transformer outperforms sequence-only under transformations
🔎 Similar Papers
No similar papers found.