🤖 AI Summary
This work addresses the challenges of resolving direct and indirect code references in decompilation by formalizing reference rewriting as a permutation-based indexing task. The authors construct the first synthetic benchmark for this problem and introduce a novel sequence-to-sequence neural architecture that integrates permutation modeling with a dedicated indexing mechanism. The proposed model significantly outperforms existing approaches in both robustness and scalability: it handles input sequences ten times longer than those processed by the strongest baseline and reduces error rates by 42% on decompiling switch statements. Ablation studies confirm the necessity of each architectural component.
📝 Abstract
Resolving and rewriting references is fundamental in programming languages. Motivated by a real-world decompilation task, we abstract reference rewriting into the problems of direct and indirect indexing by permutation. We create synthetic benchmarks for these tasks and show that well-known sequence-to-sequence machine learning architectures are struggling on these benchmarks. We introduce new sequence-to-sequence architectures for both problems. Our measurements show that our architectures outperform the baselines in both robustness and scalability: our models can handle examples that are ten times longer compared to the best baseline. We measure the impact of our architecture in the real-world task of decompiling switch statements, which has an indexing subtask. According to our measurements, the extended model decreases the error rate by 42%. Multiple ablation studies show that all components of our architectures are essential.