Score
Designs and implements software that directly executes programs written in a programming or domain-specific language, including components such as lexical analysis, parsing, AST or bytecode representation, evaluation/execution engine, runtime environment, memory and environment management, and error reporting. Responsible for interpreter architecture, correctness and performance (e.g., bytecode generation, execution strategies, integration points for JITs), and integration with tooling such as REPLs, debuggers, and foreign‑function interfaces.
Reimplementing abstract interpreters for new programming languages incurs high development costs. Method: This paper proposes an automatic abstract interpreter retargeting technique based on partial evaluation. It embeds the formal semantics of a target language into the source language of an existing abstract interpreter and achieves cross-language analyzer migration via semantics-driven specialization—requiring neither manual rewriting nor adaptation. Contribution/Results: The core innovation lies in the tight integration of semantic modeling, abstract interpretation, and partial evaluation to guarantee both semantic correctness and analysis precision of the migrated analyzer. Experimental evaluation demonstrates the method’s feasibility and effectiveness across diverse programming languages, significantly improving static analyzer reuse efficiency and construction speed while preserving soundness and precision.
This study investigates whether domain-specific languages (DSLs) enhance developers’ comprehension of data pipeline program structure. Method: A mixed-methods approach is employed—controlled experiments measure task accuracy, while structured surveys and qualitative coding analyze DSLs’ impact on domain experts’ structural awareness, accessibility, and alignment with mental models. Contribution/Results: This work provides the first empirical validation of systematic improvements in structural understanding of data pipelines afforded by DSLs. Results show statistically significant gains in comprehension accuracy (p < 0.01), driven by DSLs’ capacity to reinforce global program overviews, enforce syntactically constrained structures, and better align with users’ domain-specific mental models. Furthermore, DSLs lower the barrier to entry for programmers with limited experience, facilitate cross-tool knowledge transfer, and strengthen perception of dataflow structure.
This work investigates the impact of incorporating runtime execution information—specifically line coverage, branch coverage, variable states, and execution frequency—into large language models for code (LLMs) on automated code optimization performance. Building upon CodeT5+, we systematically design three execution-aware pretraining strategies and, for the first time, quantitatively evaluate the individual and joint contributions of these four execution dimensions to efficiency-oriented optimization within a unified framework. Experimental results reveal only marginal performance gains from execution-aware modeling; several configurations even underperform the baseline significantly—challenging the widely held implicit assumption that execution signals are inherently beneficial. This study critically questions the prevailing consensus on the efficacy of execution-augmented modeling paradigms and provides key empirical evidence and reflective insights for jointly advancing interpretability and practical utility in code LLMs.
Binary program symbolic execution suffers from semantic distortion and implementation errors introduced during intermediate representation (IR) translation. Method: This paper proposes the first instruction-level symbolic execution framework directly grounded in formal ISA semantics (Rock/Sail), bypassing conventional IR abstractions by compiling machine-readable ISA specifications into SMT-solvable symbolic semantic models and integrating them into a binary analysis platform. Contributions/Results: (1) The first end-to-end automated pipeline from formal ISA semantics to symbolic execution; (2) Demonstrated scalability on RISC-V—modeling new instructions requires only a few hours; (3) Discovered five previously unknown ISA semantic implementation bugs in angr; (4) Achieved high-fidelity branch modeling and solving capability. The framework significantly improves the accuracy, trustworthiness, and development efficiency of binary symbolic execution.
Large language models (LLMs) exhibit pervasive output formatting bias in code translation tasks—generated outputs frequently contain extraneous natural-language explanations or formatting delimiters, causing standard evaluation metrics (e.g., computation accuracy, CA) to systematically underestimate true performance. Method: We systematically evaluate 11 instruction-tuned LLMs across five programming languages and find that 26.4%–73.7% of translations require post-hoc processing to extract clean code. To address this, we propose a robust code extraction method integrating regex-based parsing with prompt engineering. Contribution/Results: Our approach achieves a 92.73% average Code Extraction Success Rate (CSR) on a multilingual alignment benchmark, substantially improving evaluation fidelity. This work is the first to quantify the impact of formatting bias and establishes a new, generalizable, and robust code extraction paradigm—providing a reproducible, standardized evaluation benchmark for LLM-based code translation.
Modular control-flow handling in abstract interpretation and supporting multiple analysis strategies—such as path- vs. flow-sensitivity, forward vs. backward directionality, and upper vs. lower approximations—traditionally relies on complex monad transformers, leading to implementation brittleness and poor composability. Method: This paper introduces the *cumulative abstract semantics* framework, the first to incorporate *scoped effects* into abstract interpretation. It decouples syntactic structure from semantic behavior via two classes of effect handlers: *syntax-resolving* and *domain-semantics-introducing*. A single syntax-driven interpreter suffices to generate diverse dynamic evaluators and static analyzers. Contribution/Results: The framework eliminates heavyweight data structures, preserving expressiveness while drastically reducing implementation complexity for multi-strategy analyses. It enhances maintainability, composability, and modularity—providing a concise, unified, and extensible theoretical and practical foundation for modular program analysis.
This work addresses the significant performance degradation of large language models (LLMs) in generating code for constraint-based domain-specific languages (DSLs), such as OCL and Alloy, and the absence of systematic evaluation methodologies. The paper introduces the first evaluation framework tailored for constraint DSL code generation, which systematically assesses LLM capabilities in translating natural language to DSL through both syntactic correctness and semantic accuracy, leveraging formal verification. Experimental comparisons across Python, OCL, and Alloy reveal that LLMs perform markedly better on general-purpose languages, that models with limited context windows struggle to jointly generate constraints and domain models, and that incorporating code repair and multi-candidate generation strategies substantially improves output quality. The framework further enables systematic analysis of prompting templates, repair mechanisms, and multi-turn generation strategies.
Traditional imperative programming relies on predefined textual instructions, requiring developers to anticipate state changes—a cognitively demanding and error-prone process. This work proposes an execution-centric, incremental approach to program construction: by directly manipulating data and recording execution traces, program behavior is transformed into an intermediate representation, with continuous alignment between program structure and execution state maintained through an explicit conceptual machine. Control structures are deterministically synthesized from observed execution behaviors—conditionals emerge from recorded comparisons, and loops are encapsulated as macros supporting nonlinear, incremental development. The system enables partial execution, refinement, and completion while preserving semantic consistency, and generates readable, correct code in multiple languages, including Python, C, C++, and Java. Empirical validation on standard algorithmic benchmarks demonstrates the method’s correctness, expressiveness, robustness, and language independence.
This work addresses the lack of effective control over privilege escalation in existing intelligent systems during dynamic code generation and execution. We propose a "controlled metaprogramming" paradigm that treats program representations as first-class values, decoupling code generation from execution through purely syntactic manipulation and structural inspection mechanisms. Specifically, we reconceptualize eval—not as a language primitive—but as a controlled effect subject to validation against policies, capabilities, and resource constraints. Leveraging formal methods, we define pure syntactic evaluation and controlled materialization judgments, implementing them in MashinTalk, a domain-specific language compiled to BEAM bytecode. Our theoretical guarantees—encompassing purity of syntactic operations, non-bypassability, and boundary preservation—are formally verified within a suite of 454 machine-checked theorems in Rocq.
This work addresses the prevalent issue in large language models (LLMs) of introducing control-flow, type, or I/O errors during code translation due to neglect of program intent. To mitigate this, the paper proposes the first systematic use of a language-agnostic, structured intermediate specification that preserves semantic fidelity through an intermediate representation, structured generation, and automated test-based validation. Evaluated on the Avatar and CodeNet datasets with five state-of-the-art LLMs, the approach significantly improves translation accuracy, raising the micro-averaged accuracy from 67.7% to 78.5%. It completely eliminates lexical errors and substantially reduces errors related to structure, declarations, and runtime dependencies.