Score
Designs and builds methods, tools, or analyses that examine program semantics to identify and extract the essential parts of code and produce minimized programs that preserve intended behavior. This work formulates reduction hypotheses (optionally using LLM-driven heuristics), applies semantics-preserving transformations, and validates reductions through execution-based checks or other semantic verification.
Large language models (LLMs) exhibit limited understanding of compiler-level semantics-preserving program transformations—e.g., copy propagation and constant folding—critical for reliable code reasoning. Method: We propose a formal-verification–based empirical evaluation framework for semantic equivalence judgment, leveraging LLVM and other compiler toolchains to automatically generate robust test cases and self-supervised training signals. Contribution/Results: Experiments reveal high failure rates: 41% without context and 29% even with simple generic context—exposing fundamental blind spots in deep code semantic modeling. To address this, we introduce the first LLM–compiler co-enhanced training paradigm, wherein compiler-generated semantic equivalence pairs explicitly reinforce model robustness. This work establishes a rigorous methodology for quantitatively assessing code understanding capabilities and provides a scalable, tool-integrated pathway toward trustworthy code AI.
This work proposes a novel paradigm for program reduction based on a dual-agent large language model framework, reframing reduction as an autonomous reasoning task. Unlike conventional approaches that rely on predefined rules and lack semantic understanding or learning capabilities, the proposed method employs one agent to analyze program semantics and generate reduction hypotheses, while a second agent iteratively refines these hypotheses through execution feedback. The system further leverages knowledge distillation to extract transferable strategies, enabling continual self-improvement. By uniquely integrating semantic awareness with experience-driven learning, the approach significantly outperforms state-of-the-art tools across 90 benchmarks spanning three programming languages, consistently producing smaller and more efficient reduced programs.
This work investigates the capability of large language models (LLMs) in program semantic understanding and formal specification synthesis for correctness verification. To this end, we introduce FormalBench—the first comprehensive benchmark tailored to formal specification inference—covering core challenges such as loop reasoning and semantics-preserving transformations. Experimental results show that while LLMs perform well on simple control-flow structures, their robustness degrades significantly on complex loops and semantic equivalence transformations. Building on these findings, we propose a self-healing prompting strategy that iteratively validates and refines generated specifications, improving synthesis success rate by 25%. Our study provides the first systematic characterization of LLMs’ capabilities and limitations in program semantic reasoning. Moreover, it delivers a reproducible evaluation framework and effective intervention techniques to advance LLMs’ formal reasoning capacity.
Compiler testing often involves large-scale programs, which significantly hinders the efficiency of bug reproduction and debugging. To address this challenge, this work proposes SimP, a novel framework that introduces large language models (LLMs) into program reduction for the first time. SimP synergistically combines rule-based syntactic pruning with LLM-driven semantic guidance, employing customized prompts to enable multi-stage collaborative optimization. The approach achieves substantially improved reduction efficiency while maintaining comparable reduction quality to existing methods. Notably, the computational overhead and economic cost associated with LLM invocations are negligible, making SimP both high-performing and practical for real-world compiler testing scenarios.
Existing code understanding evaluation frameworks are limited to single-input reasoning tasks and narrow execution trace coverage, failing to assess large language models’ (LLMs) deep semantic comprehension of programs. Method: We propose the first black-box evaluation framework grounded in formal program specifications, centered on specifications that provably cover all execution traces. It comprises four progressively sophisticated specification-understanding tasks—from basic to advanced—and introduces counterfactual perturbation generation alongside contrastive evaluation to rigorously test model robustness to semantics-preserving transformations. Contribution/Results: Extensive evaluation across six state-of-the-art code LLMs reveals pervasive deficiencies in specification understanding and heightened sensitivity to semantically invariant perturbations—indicating fundamental flaws in their underlying semantic representations.
This work proposes a hybrid concrete-symbolic interpretation method to efficiently verify semantic equivalence between original and optimized programs in MLIR, ensuring the correctness of optimization transformations. The approach supports diverse syntactic, scheduling, and memory representations and theoretically achieves linear-time complexity for equivalence checking. Building upon this method, the authors develop a formal verifier for a subset of MLIR and successfully apply it to the AMD MLIR-AIR and MLIR-AIE toolchains as well as the standard mlir-opt infrastructure. Evaluation across hundreds of benchmark variants demonstrates the verifier’s effectiveness in validating optimization pipelines, significantly enhancing the reliability of compiler optimizations within the MLIR ecosystem.
This work addresses the challenge of statically verifying semantic consistency between natural language business requirements and their code implementations. It proposes a two-stage, runtime-free approach: first leveraging large language models to extract structured rules from requirements while identifying ambiguous or contradictory statements, and then performing static code auditing based on this intermediate representation. By integrating natural language processing with static analysis, the method mitigates hallucination and context loss in large models through rule structuring, enabling requirement-aware early validation. Evaluated on an automotive cybersecurity case study, the approach successfully detects semantic deviations, offers a novel solution to the test oracle problem, and significantly enhances left-shifted verification capabilities.
Traditional symbolic execution struggles in complex systems like Android due to path explosion, difficulties in function modeling, and semantic loss, hindering efficient extraction of precise constraints. This work proposes RECON, a novel framework that integrates large language models (LLMs) into backward constraint analysis: starting from a target method and tracing back to entry points, it extracts method-level control-flow constraints and leverages an LLM to translate bytecode conditions into interpretable semantic specifications. While preserving logical equivalence, this approach substantially enhances both analysis efficiency and interpretability. Experimental results demonstrate that RECON achieves 100% success across 78 Android scenarios, operating 5.8× faster than conventional symbolic execution, and generates semantic constraints triggering dangerous APIs with 84% success on 100 malicious samples.
Existing code translation evaluation metrics, such as BLEU, rely solely on syntactic similarity and fail to capture semantic correctness. This work introduces, for the first time, the principles of compiler testing into the evaluation of large language models (LLMs) for code translation, proposing a semantic equivalence verification framework grounded in execution consistency. The study defines "semantic accuracy" as the core evaluation metric and develops an LLM-based decompiler that integrates execution trace comparison, semantic equivalence validation, and LLM fine-tuning. Experimental results demonstrate that the proposed LLM decompiler significantly outperforms heuristic baselines, while revealing a negligible correlation between BLEU scores and semantic accuracy (r = –0.127 to 0.354), thereby confirming BLEU’s inadequacy as a proxy for functional correctness in code translation tasks.
This study addresses the lack of systematic evaluation of large language models’ (LLMs) ability to automatically generate formal preconditions and postconditions from natural language specifications. The authors introduce a new benchmark dataset comprising 40 tasks and conduct the first comprehensive assessment of 24 state-of-the-art LLMs on this task. They propose a fine-grained validation approach that integrates automatically generated test cases and employs multi-dimensional metrics—including accept@1 and refined correctness analysis—to evaluate the generated formal specifications. Their findings reveal that LLMs perform better at formalizing preconditions than postconditions and that closed-source models generally outperform open-source ones. Furthermore, the proposed automated testing mechanism effectively identifies incorrect formalizations erroneously deemed correct by conventional evaluation methods, substantially enhancing assessment reliability.