compilers

Design, implement, and evaluate program translation systems that take source code and produce intermediate representations or target code, encompassing lexing, parsing, semantic analysis, type checking, optimization passes, and code generation. Build and analyze compiler transformations and passes for correctness, performance, and portability, and integrate them into toolchains and runtimes.

compilers

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$217K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Large language models (LLMs) exhibit pervasive output formatting bias in code translation tasks—generated outputs frequently contain extraneous natural-language explanations or formatting delimiters, causing standard evaluation metrics (e.g., computation accuracy, CA) to systematically underestimate true performance. Method: We systematically evaluate 11 instruction-tuned LLMs across five programming languages and find that 26.4%–73.7% of translations require post-hoc processing to extract clean code. To address this, we propose a robust code extraction method integrating regex-based parsing with prompt engineering. Contribution/Results: Our approach achieves a 92.73% average Code Extraction Success Rate (CSR) on a multilingual alignment benchmark, substantially improving evaluation fidelity. This work is the first to quantify the impact of formatting bias and establishes a new, generalizable, and robust code extraction paradigm—providing a reproducible, standardized evaluation benchmark for LLM-based code translation.

Evaluating LLM code translation suffers from output format biasesNon-code elements in outputs interfere with performance assessment metricsProposing methods to extract source code for reliable model evaluation

Existing code translation evaluation metrics, such as BLEU, rely solely on syntactic similarity and fail to capture semantic correctness. This work introduces, for the first time, the principles of compiler testing into the evaluation of large language models (LLMs) for code translation, proposing a semantic equivalence verification framework grounded in execution consistency. The study defines "semantic accuracy" as the core evaluation metric and develops an LLM-based decompiler that integrates execution trace comparison, semantic equivalence validation, and LLM fine-tuning. Experimental results demonstrate that the proposed LLM decompiler significantly outperforms heuristic baselines, while revealing a negligible correlation between BLEU scores and semantic accuracy (r = –0.127 to 0.354), thereby confirming BLEU’s inadequacy as a proxy for functional correctness in code translation tasks.

BLEUcode translationprogram semantics

Repository-Level Compositional Code Translation and Validation

Oct 31, 2024
AR
Ali Reza Ibrahimzada
🏛️ University of Illinois Urbana-Champaign | Indian Institute of Science | Cornell University | IBM Research

This paper addresses the challenge of ensuring functional consistency in cross-language code translation by proposing AlphaTrans, an end-to-end neural-symbolic fusion framework designed for real-world open-source projects. Methodologically, it introduces a novel *inverse call-order divide-and-conquer translation strategy*, integrating static program analysis with call-graph-driven code slicing to jointly translate source code and corresponding tests; it further establishes a three-tier automated verification mechanism—syntactic checking, unit test execution, and assertion-level output comparison. Contributions include robust support for industrial-scale projects featuring complex dependencies, custom types, and language-specific constructs. Evaluated on 10 real-world projects (836 classes, 8,575 methods, 2,719 tests), AlphaTrans achieves 96.4% syntactic correctness and 25.14% functional pass rate, with an average translation time of 34 hours per project. Developers require only 20.1 hours on average—guided by diagnostic reports—to achieve full test-suite passing.

Automate repository-level code translation across languages.Ensure functionality preservation post-translation via validation.Scale translation to real-world projects with dependencies.

Traditional compilers face limitations in development accessibility, optimization capabilities, and application scope. This work proposes the first multidimensional classification framework for large language model (LLM)-driven compilation, offering a systematic survey of existing research through four analytical dimensions: design philosophy, methodology, level of code abstraction, and task type. The study identifies three core design paradigms—Selector, Translator, and Generator—and highlights three transformative directions: democratizing compiler development, discovering novel optimization strategies, and expanding functional boundaries. It further argues that hybrid systems represent a critical pathway forward and provides a technical roadmap for building correct, scalable, and intelligent compilation tools.

compiler correctnesscompiler optimizationhybrid systems

Latest Papers

What's happening recently
View more

This work addresses the interoperability challenge between GCC and LLVM compiler intermediate representations (IRs), which stems from their semantic and structural differences. To bridge this gap, the authors propose IRIS-14B, the first large language model specifically designed for IR-to-IR translation. Built upon a 14-billion-parameter Transformer architecture, IRIS-14B leverages supervised fine-tuning to learn the mapping between GIMPLE and LLVM IR derived from the same C source code, enabling high-fidelity automatic translation. Experimental results demonstrate that IRIS-14B substantially outperforms existing open-source large models on real-world C programs and competitive programming tasks, achieving up to a 44-percentage-point improvement in accuracy. This study provides the first empirical validation of large language models as effective and feasible interoperability layers within neuro-symbolic hybrid compilation frameworks.

Compiler Intermediate RepresentationCross-toolchain InteroperabilityGIMPLE

This work addresses the persistent gap in runtime efficiency between code translations generated by large language models (LLMs) and those written by humans—a limitation that is difficult to mitigate through prompt engineering alone. To bridge this gap, the authors propose SwiftTrans, a novel framework that first generates diverse translation candidates through multi-perspective exploration and then selects the optimal solution using a discrepancy-aware selection mechanism. The framework further incorporates hierarchical and ordinal guidance strategies to enhance performance. Notably, this study is the first to systematically balance both functional correctness and runtime efficiency in LLM-based code translation. The authors introduce SwiftBench, a new benchmark tailored for this dual objective, and demonstrate that SwiftTrans achieves consistent and significant improvements over existing methods across CodeNet, F2SBench, and SwiftBench.

code translationfunctional correctnesslarge language models

Current evaluations of code translation often misattribute model failures to incorrect outputs when, in fact, the errors stem from improper compilation flags, library linking issues, or environmental misconfigurations—thereby obscuring true model performance. This work presents the first systematic identification and categorization of such model-agnostic “pseudo-failures.” Analyzing 6,164 code translation samples generated by GPT-4o, DeepSeek-Coder, and Magicoder across five languages (C, C++, Java, Python, and Go) on the Avatar, CodeNet, and EvalPlus benchmarks, the study quantifies the substantial impact of evaluation configuration flaws on reported results. It reveals that a significant portion of alleged failures are actually due to environmental errors rather than logical inaccuracies. The findings advocate shifting the focus of code translation evaluation from mere logical correctness toward end-to-end reliability and call for transparent, configuration-aware evaluation standards.

code translationevaluation reliabilityfalse failures

This work addresses the limitations of current large language models in code translation, which rely heavily on superficial statistical patterns and lack deep program semantic understanding, particularly in real-world scenarios where high-quality semantic supervision is often unavailable. To overcome this, the authors propose Multisage, a novel framework that automatically constructs multi-dimensional structured semantic representations—such as data-flow graphs, type constraints, and API usage—from source code and generates diverse semantic augmentation signals, including natural language summaries, test cases, and API descriptions. The framework incorporates a self-calibration mechanism through semantic-preserving mutations and cross-semantic consistency verification, eliminating the need for external annotations. Evaluated on the HumanEval-X benchmark, Multisage improves translation success rates by up to 2.22× over state-of-the-art prompting, fine-tuning, and chain-of-thought approaches, with especially pronounced gains on smaller models.

code translationlarge language modelsprogram semantics

Existing approaches to automatic C-to-Rust translation suffer from semantic loss and structural distortion due to preprocessing steps that eliminate macros. This work proposes a synergistic “relay” translation strategy that combines formal methods with large language models (LLMs): a formal translator, MerC, first handles reducible macros according to rigorously defined rules, while an LLM subsequently processes more complex macro constructs. We establish the first formal specification for macro translation and introduce MacroBench, a dedicated benchmark for evaluating macro translation quality. Experimental results demonstrate that MerC correctly translates 50% of macros on MacroBench, whereas LLMs exhibit broader coverage but error rates ranging from 8% to 28%. The proposed collaborative approach translates 51% more test cases on average and reduces failure rates by 32%, substantially improving both accuracy and completeness in macro translation.

C to Rustcode translationformal specification