Score
Designs and implements systems that use large language models to synthesize structured or unstructured inputs and test cases—including adversarial, size- or complexity-controlled, and API- or language-targeted examples—by prompting, conditioning, or otherwise steering model generation. Builds tooling to guide generation with auxiliary signals (e.g., graphs), adapt outputs to target languages/APIs, and integrate execution and validation to probe worst-case and failure behaviors.
Large language models (LLMs) exhibit limited autonomous capability in solving open-ended problems, primarily due to overreliance on explicit algorithms and static knowledge. Method: We propose a novel end-to-end paradigm—spanning problem framing, solution exploration, implementation generation, and strategy assessment—that integrates prompt engineering, retrieval-augmented generation (RAG), and reinforcement learning from human feedback (RLHF). This synergy enhances LLMs’ proficiency in feature composition, dynamic anomaly response, and high-level strategy evaluation. Contribution/Results: We present the first systematic taxonomy of paradigm evolution for LLM-based implementation generation, identify critical technical bottlenecks, and establish a theoretical framework and technology roadmap for autonomous problem solving. Our work advances the development of LLM-driven general-purpose agents by enabling more robust, adaptive, and self-assessing reasoning capabilities.
Current large language models (LLMs) ensure syntactic and constraint validity in structured generation but suffer from severely limited output diversity. To address this, we propose an automaton-guided generation mechanism that leverages historical state-transition trajectories—extracted during structured decoding—to dynamically steer the model toward under-explored structural patterns. By tightly integrating automata theory with LLM decoding, our method enhances structural and semantic diversity without compromising validity or inference efficiency. Experimental evaluation on open-source library test-case generation demonstrates a 27.4% improvement in diversity metrics—including structural coverage and semantic dissimilarity—while maintaining a 98.6% compliance rate with syntax and domain constraints. This confirms the method’s effectiveness and practical applicability for diverse, valid structured generation.
This work investigates large language models’ (LLMs) iterative input-output (I/O) reasoning capability in example-driven code generation—specifically, inferring functional intent and generalizing correct code from sparse, ambiguous, or incomplete I/O examples. To this end, the authors introduce the first comprehensive evaluation framework for this task, proposing a “fitting → generalization” two-phase assessment paradigm and releasing a new benchmark comprising 168 functionally diverse tasks. Experiments span six state-of-the-art models—including GPT-4, Claude, and Llama variants—revealing three key findings: (1) model performance drops by over 60% compared to natural-language instruction-based code generation; (2) 95% of successful generations occur on the first prompt iteration, with subsequent iterations rarely enabling effective correction; and (3) these results expose fundamental bottlenecks in LLMs’ efficiency in leveraging sparse examples and their capacity for functional generalization.
This work addresses black-box program behavior modeling by proposing a reversible, differentiable, and constraint-aware, grammar-driven neural modeling framework. Methodologically, it generates input-output (I/O) pairs from formal grammars of input and output languages, and employs a lightweight (<6.3M-parameter) sequence-to-sequence model to cast program I/O mapping as a bidirectional neural machine translation task—enabling both forward prediction and backward inference—while supporting fine-grained behavioral constraints and fault- or coverage-guided input synthesis. Its key contribution is the first end-to-end joint modeling of reversibility, differentiability, and syntactic consistency in program behavior models. Evaluated on structured tasks such as Markdown and HTML generation, the framework achieves 95.4% accuracy and a BLEU score of 0.98±0.04, significantly outperforming existing irreversible or syntax-agnostic approaches.
This work addresses the underexplored challenge of automatically generating formal specifications—expressed in first-order logic—from software comments and documentation using large language models (LLMs). Method: We conduct the first systematic evaluation of 13 state-of-the-art LLMs (e.g., Codex, Llama, PaLM) against traditional approaches on three public benchmarks under few-shot settings. We introduce a cross-model failure diagnosis framework, establish a reproducible evaluation benchmark, and propose a taxonomy of failure modes. Contribution/Results: Experiments reveal that certain LLMs achieve performance comparable to or exceeding traditional tools in specific scenarios; however, semantic abstraction, context sensitivity, and logical rigor remain critical bottlenecks. Our analysis uncovers complementary strengths between LLMs and classical methods, providing empirical foundations and concrete directions for advancing LLM-augmented formal methods. The benchmark, taxonomy, and diagnostic framework are publicly released to support reproducible research.
This study addresses the lack of a systematic review on the application of large language models (LLMs) in software engineering documentation and modeling tasks. Through a comprehensive literature survey, it establishes a multi-dimensional taxonomy that categorizes existing research by task type, offering an in-depth analysis of key technical approaches—including prompt engineering, natural language understanding, and structured language processing. The work further synthesizes the distribution of tasks, evaluation metrics, human assessment methodologies, and commonly used datasets across major conferences in the field. By systematically mapping the research landscape and identifying prevailing technical trends, this paper provides a thorough reference and strategic guidance for future investigations at the intersection of LLMs and software engineering.
This study investigates the feasibility of leveraging low-cost, deployable open-source large language models (ranging from 0.5B to 32B parameters) to automatically generate domain-specific language (DSL) representations of UI and data models directly from natural language prompts, without fine-tuning and using only few-shot prompting. It presents the first systematic evaluation of small-scale open-source models on the task of generating multiple, interrelated DSL artifacts. Through a combination of DSL grammar parsing, automated validation, and expert assessment, the work examines model performance in terms of syntactic correctness, semantic completeness, and cross-model referential consistency. Experimental results demonstrate that compact models—such as gemma3:12b and mistral:7b-instruct—achieve generation quality comparable to or even rivaling that of significantly larger models, highlighting their practical viability and cost-effectiveness for model-driven engineering applications.
This study investigates whether large language models (LLMs) underperform in generating domain-specific languages—such as AMPL for algebraic modeling—compared to general-purpose programming languages like Python, specifically within mathematical optimization contexts. To address this, the authors propose EXEOS, a method that leverages LLMs to translate natural language descriptions into either AMPL or Python code, augmented with a solver-feedback-driven iterative refinement mechanism to enhance executability and correctness. The first systematic comparison of its kind demonstrates that, across public benchmarks and real-world Kinaxis supply chain cases, LLM-generated AMPL code matches or even surpasses Python in quality. These findings affirm the competitiveness of domain-specific languages in specialized optimization tasks and highlight the critical role of solver-in-the-loop iterative refinement in improving the generation of formal specifications.
This study investigates the use of large language models (LLMs) to automatically translate neutral graph representations of fluid systems into high-quality, functionally correct code executable in mainstream simulation environments such as WNTR and Modelica. The authors systematically evaluate ten state-of-the-art LLMs combined with six prompting strategies across multiple benchmark scenarios, assessing generated code through software quality metrics and simulation fidelity. This work presents the first systematic comparison in the domain of fluid system modeling that examines how different LLMs and prompt engineering techniques influence both syntactic correctness and functional fidelity of generated simulation code, offering empirical guidance for model-driven code generation. Experimental results demonstrate that optimal configurations can produce syntactically valid code; however, a significant gap remains in achieving high simulation fidelity, highlighting key directions for future improvement.
This work addresses the challenges of applying large language models (LLMs) in modeling and simulation (M&S), where suboptimal prompt design, improper hyperparameter configuration, or inadequate data handling often lead to performance degradation, information loss, and non-deterministic behavior. For the first time, this study systematically identifies latent pitfalls specific to LLM deployment in M&S and proposes a principled framework centered on rigorous design and empirical evaluation. The framework encompasses key techniques including prompt engineering, retrieval-augmented generation (RAG), low-rank adaptation (LoRA), temperature control, and context management. By offering a structured set of practical guidelines, this research enables practitioners to critically assess the suitability and implementation strategies of LLMs in M&S contexts, thereby substantially enhancing their effectiveness and reliability.