Translation Tag Team: Formal Rules and LLMs Translate More Macros Together than Apart

πŸ“… 2026-08-06
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing approaches to automatic C-to-Rust translation suffer from semantic loss and structural distortion due to preprocessing steps that eliminate macros. This work proposes a synergistic β€œrelay” translation strategy that combines formal methods with large language models (LLMs): a formal translator, MerC, first handles reducible macros according to rigorously defined rules, while an LLM subsequently processes more complex macro constructs. We establish the first formal specification for macro translation and introduce MacroBench, a dedicated benchmark for evaluating macro translation quality. Experimental results demonstrate that MerC correctly translates 50% of macros on MacroBench, whereas LLMs exhibit broader coverage but error rates ranging from 8% to 28%. The proposed collaborative approach translates 51% more test cases on average and reduces failure rates by 32%, substantially improving both accuracy and completeness in macro translation.
πŸ“ Abstract
Modern critical software infrastructure is largely written in C. Since C lacks memory safety, researchers are investigating automatic translation of C to safer languages like Rust. But real-world C software consists of more than just C code, often using named code fragments called macros which are not part of the C language proper. State-of-the-art techniques avoid translating macros by preprocessing C code first before translating it. But this approach produces translations that are dissimilar to the original C code, because preprocessing inlines all macro definitions. To preserve macro usage in translated code, we study the language features that macros and C share and distill them into the first formally-specified translator, MerC. To evaluate MerC, we introduce the first macro translation benchmark, MacroBench, with test cases based on macros randomly sampled from real-world C programs. We find that MerC supports 50% of MacroBench's macro test cases. We also use MacroBench to evaluate how effective large language models (LLMs) are at performing the previously-unstudied task of macro translation. LLMs translate 22% to 77% more of MacroBench than MerC, but with 8% and 28% of these translations being incorrect translations requiring additional validation by developers. In contrast, MerC only produces correct translations. Our key insight is that running MerC first then using LLMs on the remainder reaps greater benefits than using either technique alone. This tag team approach has an average failure rate 32% lower than that of LLMs, while also translating an average of 51% more test cases than MerC.
Problem

Research questions and friction points this paper is trying to address.

macro translation
C to Rust
memory safety
code translation
formal specification
Innovation

Methods, ideas, or system contributions that make the work stand out.

macro translation
formal translation
large language models
C to Rust
translation benchmark
πŸ”Ž Similar Papers
No similar papers found.
B
Brent Pappas
University of Central Florida
J
Joseph Zalusky
University of Central Florida
Z
Zachary Burkett
University of Central Florida
P
Paul Gazzillo
University of Central Florida