Neural architectures for resolving references in program code

📅 2026-04-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of resolving direct and indirect code references in decompilation by formalizing reference rewriting as a permutation-based indexing task. The authors construct the first synthetic benchmark for this problem and introduce a novel sequence-to-sequence neural architecture that integrates permutation modeling with a dedicated indexing mechanism. The proposed model significantly outperforms existing approaches in both robustness and scalability: it handles input sequences ten times longer than those processed by the strongest baseline and reduces error rates by 42% on decompiling switch statements. Ablation studies confirm the necessity of each architectural component.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageMachine Learning: Deep Neural Architectures and Foundation ModelsCognitive Modeling & Cognitive Systems: Neural Spike Coding

Application Category

Graph Algorithms and Modeling for the Web: Querying, indexing, and retrieval in Web-related graphsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
Resolving and rewriting references is fundamental in programming languages. Motivated by a real-world decompilation task, we abstract reference rewriting into the problems of direct and indirect indexing by permutation. We create synthetic benchmarks for these tasks and show that well-known sequence-to-sequence machine learning architectures are struggling on these benchmarks. We introduce new sequence-to-sequence architectures for both problems. Our measurements show that our architectures outperform the baselines in both robustness and scalability: our models can handle examples that are ten times longer compared to the best baseline. We measure the impact of our architecture in the real-world task of decompiling switch statements, which has an indexing subtask. According to our measurements, the extended model decreases the error rate by 42%. Multiple ablation studies show that all components of our architectures are essential.
Problem

Research questions and friction points this paper is trying to address.

reference resolution
decompilation
indexing
program code
switch statements
Innovation

Methods, ideas, or system contributions that make the work stand out.

reference resolution
sequence-to-sequence architecture
permutation-based indexing
decompilation
synthetic benchmark
G
Gergő Szalay
Faculty of Informatics, ELTE Eötvös Loránd University, Budapest, H-1117
G
Gergely Zsolt Kovács
Faculty of Informatics, ELTE Eötvös Loránd University, Budapest, H-1117
S
Sándor Teleki
Faculty of Informatics, ELTE Eötvös Loránd University, Budapest, H-1117
Balázs Pintér
Balázs Pintér
Eötvös Loránd University
Machine LearningNatural Language Processing
T
Tibor Gregorics
Faculty of Informatics, ELTE Eötvös Loránd University, Budapest, H-1117