🤖 AI Summary
This study addresses the lagging AArch64 software ecosystem and the limitations of existing binary translation approaches, namely high runtime overhead in dynamic systems and low reliability in static ones. To overcome these challenges, this work proposes a large language model (LLM)-driven paradigm for static assembly-to-assembly translation. It pioneers an LLM-based static translation architecture that generates native code through precompilation, thereby eliminating dependencies on runtime frameworks. Furthermore, it introduces an efficient semantic verification mechanism based on simplified fragments to ensure translation correctness. Experimental results demonstrate that the proposed system significantly outperforms mainstream open-source solutions and the industrial-grade tool ExaGear, achieving near-native execution efficiency alongside high reliability.
📝 Abstract
While AArch64 CPUs are becoming strong market contenders, their software ecosystem lags behind the mature x86-64 environment, hindering the adoption of the new architectures and impacting user experience. Binary translation bridges this divide by converting binary code from one architecture (e.g., x86-64) to run on another (e.g., AArch64), allowing legacy software to benefit from modern hardware's performance and energy efficiency advantages. Current translation methods are typically either dynamic, which adds significant runtime overhead, or static, which struggles with reliability due to the inherent complexities of binary analysis. This paper introduces a new static, assembly-to-assembly translation paradigm that transforms binary code ahead of execution, generating portable, efficient nativelike binaries that run on AArch64 devices without runtime frameworks. Benefiting from recent breakthroughs in large language models (LLMs), we provide a practical and automated translation engine that produces high-quality code with minimal human intervention. To ensure correctness, we introduce a crucial verification step, where we split the assembly code into simplified snippets, enabling efficient and scalable semantic verification. Our evaluation shows that this approach significantly outperforms existing open-source solutions with a large margin, producing binaries with near-native performance. Furthermore, it shows substantial improvements over the leading industrial translator, ExaGear, illuminating a promising new direction for cross-architecture binary translation research.