llm-assisted input synthesis

Designs and implements systems that use large language models to synthesize structured or unstructured inputs and test cases—including adversarial, size- or complexity-controlled, and API- or language-targeted examples—by prompting, conditioning, or otherwise steering model generation. Builds tooling to guide generation with auxiliary signals (e.g., graphs), adapt outputs to target languages/APIs, and integrate execution and validation to probe worst-case and failure behaviors.

llm-assistedinputsynthesis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.65
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$203K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Large language models (LLMs) exhibit limited autonomous capability in solving open-ended problems, primarily due to overreliance on explicit algorithms and static knowledge. Method: We propose a novel end-to-end paradigm—spanning problem framing, solution exploration, implementation generation, and strategy assessment—that integrates prompt engineering, retrieval-augmented generation (RAG), and reinforcement learning from human feedback (RLHF). This synergy enhances LLMs’ proficiency in feature composition, dynamic anomaly response, and high-level strategy evaluation. Contribution/Results: We present the first systematic taxonomy of paradigm evolution for LLM-based implementation generation, identify critical technical bottlenecks, and establish a theoretical framework and technology roadmap for autonomous problem solving. Our work advances the development of LLM-driven general-purpose agents by enabling more robust, adaptive, and self-assessing reasoning capabilities.

Complex Problem SolvingLarge Language ModelsOpen-ended Questions

Current large language models (LLMs) ensure syntactic and constraint validity in structured generation but suffer from severely limited output diversity. To address this, we propose an automaton-guided generation mechanism that leverages historical state-transition trajectories—extracted during structured decoding—to dynamically steer the model toward under-explored structural patterns. By tightly integrating automata theory with LLM decoding, our method enhances structural and semantic diversity without compromising validity or inference efficiency. Experimental evaluation on open-source library test-case generation demonstrates a 27.4% improvement in diversity metrics—including structural coverage and semantic dissimilarity—while maintaining a 98.6% compliance rate with syntax and domain constraints. This confirms the method’s effectiveness and practical applicability for diverse, valid structured generation.

Enhancing diversity in automaton-based structured generation for LLMsOvercoming limited output diversity in structured generation methodsSteering LLMs toward novel structural patterns using automata history

This work investigates large language models’ (LLMs) iterative input-output (I/O) reasoning capability in example-driven code generation—specifically, inferring functional intent and generalizing correct code from sparse, ambiguous, or incomplete I/O examples. To this end, the authors introduce the first comprehensive evaluation framework for this task, proposing a “fitting → generalization” two-phase assessment paradigm and releasing a new benchmark comprising 168 functionally diverse tasks. Experiments span six state-of-the-art models—including GPT-4, Claude, and Llama variants—revealing three key findings: (1) model performance drops by over 60% compared to natural-language instruction-based code generation; (2) 95% of successful generations occur on the first prompt iteration, with subsequent iterations rarely enabling effective correction; and (3) these results expose fundamental bottlenecks in LLMs’ efficiency in leveraging sparse examples and their capacity for functional generalization.

Assessing LLMs' ability to infer functionalities from I/O examplesEvaluating LLMs on iterative example-based code generationExploring prompt optimization for improved code generation performance

Learning Program Behavioral Models from Synthesized Input-Output Pairs

Jul 11, 2024
TM
Tural Mammadov
🏛️ CISPA Helmholtz Center for Information Security | Saarland University

This work addresses black-box program behavior modeling by proposing a reversible, differentiable, and constraint-aware, grammar-driven neural modeling framework. Methodologically, it generates input-output (I/O) pairs from formal grammars of input and output languages, and employs a lightweight (<6.3M-parameter) sequence-to-sequence model to cast program I/O mapping as a bidirectional neural machine translation task—enabling both forward prediction and backward inference—while supporting fine-grained behavioral constraints and fault- or coverage-guided input synthesis. Its key contribution is the first end-to-end joint modeling of reversibility, differentiability, and syntactic consistency in program behavior models. Evaluated on structured tasks such as Markdown and HTML generation, the framework achieves 95.4% accuracy and a BLEU score of 0.98±0.04, significantly outperforming existing irreversible or syntax-agnostic approaches.

Assists in program understanding and maintenance through synthesis.Learns program behavior models from input-output pairs.Predicts outputs and inputs using neural machine translation.

How Effective are Large Language Models in Generating Software Specifications?

Jun 06, 2023
DX
Danning Xie
🏛️ Purdue University | UNIST | IBM

This work addresses the underexplored challenge of automatically generating formal specifications—expressed in first-order logic—from software comments and documentation using large language models (LLMs). Method: We conduct the first systematic evaluation of 13 state-of-the-art LLMs (e.g., Codex, Llama, PaLM) against traditional approaches on three public benchmarks under few-shot settings. We introduce a cross-model failure diagnosis framework, establish a reproducible evaluation benchmark, and propose a taxonomy of failure modes. Contribution/Results: Experiments reveal that certain LLMs achieve performance comparable to or exceeding traditional tools in specific scenarios; however, semantic abstraction, context sensitivity, and logical rigor remain critical bottlenecks. Our analysis uncovers complementary strengths between LLMs and classical methods, providing empirical foundations and concrete directions for advancing LLM-augmented formal methods. The benchmark, taxonomy, and diagnostic framework are publicly released to support reproducible research.

Compare LLMs with traditional specification extraction methodsDiagnose failures in LLMs and traditional methodsEvaluate LLMs in generating software specifications

Latest Papers

What's happening recently
View more

This study addresses the lack of a systematic review on the application of large language models (LLMs) in software engineering documentation and modeling tasks. Through a comprehensive literature survey, it establishes a multi-dimensional taxonomy that categorizes existing research by task type, offering an in-depth analysis of key technical approaches—including prompt engineering, natural language understanding, and structured language processing. The work further synthesizes the distribution of tasks, evaluation metrics, human assessment methodologies, and commonly used datasets across major conferences in the field. By systematically mapping the research landscape and identifying prevailing technical trends, this paper provides a thorough reference and strategic guidance for future investigations at the intersection of LLMs and software engineering.

Generative AILarge Language ModelsSoftware Documentation

This study investigates the feasibility of leveraging low-cost, deployable open-source large language models (ranging from 0.5B to 32B parameters) to automatically generate domain-specific language (DSL) representations of UI and data models directly from natural language prompts, without fine-tuning and using only few-shot prompting. It presents the first systematic evaluation of small-scale open-source models on the task of generating multiple, interrelated DSL artifacts. Through a combination of DSL grammar parsing, automated validation, and expert assessment, the work examines model performance in terms of syntactic correctness, semantic completeness, and cross-model referential consistency. Experimental results demonstrate that compact models—such as gemma3:12b and mistral:7b-instruct—achieve generation quality comparable to or even rivaling that of significantly larger models, highlighting their practical viability and cost-effectiveness for model-driven engineering applications.

Domain Specific LanguagesGrammar-Based GenerationLarge Language Models

This study investigates whether large language models (LLMs) underperform in generating domain-specific languages—such as AMPL for algebraic modeling—compared to general-purpose programming languages like Python, specifically within mathematical optimization contexts. To address this, the authors propose EXEOS, a method that leverages LLMs to translate natural language descriptions into either AMPL or Python code, augmented with a solver-feedback-driven iterative refinement mechanism to enhance executability and correctness. The first systematic comparison of its kind demonstrates that, across public benchmarks and real-world Kinaxis supply chain cases, LLM-generated AMPL code matches or even surpasses Python in quality. These findings affirm the competitiveness of domain-specific languages in specialized optimization tasks and highlight the critical role of solver-in-the-loop iterative refinement in improving the generation of formal specifications.

Algebraic SpecificationsCode GenerationDomain-Specific Language

This study investigates the use of large language models (LLMs) to automatically translate neutral graph representations of fluid systems into high-quality, functionally correct code executable in mainstream simulation environments such as WNTR and Modelica. The authors systematically evaluate ten state-of-the-art LLMs combined with six prompting strategies across multiple benchmark scenarios, assessing generated code through software quality metrics and simulation fidelity. This work presents the first systematic comparison in the domain of fluid system modeling that examines how different LLMs and prompt engineering techniques influence both syntactic correctness and functional fidelity of generated simulation code, offering empirical guidance for model-driven code generation. Experimental results demonstrate that optimal configurations can produce syntactically valid code; however, a significant gap remains in achieving high simulation fidelity, highlighting key directions for future improvement.

code synthesisfluid systemslarge language models

This work addresses the challenges of applying large language models (LLMs) in modeling and simulation (M&S), where suboptimal prompt design, improper hyperparameter configuration, or inadequate data handling often lead to performance degradation, information loss, and non-deterministic behavior. For the first time, this study systematically identifies latent pitfalls specific to LLM deployment in M&S and proposes a principled framework centered on rigorous design and empirical evaluation. The framework encompasses key techniques including prompt engineering, retrieval-augmented generation (RAG), low-rank adaptation (LoRA), temperature control, and context management. By offering a structured set of practical guidelines, this research enables practitioners to critically assess the suitability and implementation strategies of LLMs in M&S contexts, thereby substantially enhancing their effectiveness and reliability.

Hyper-parameter TuningLarge Language ModelsModeling and Simulation

Hot Scholars

PN

Preslav Nakov

Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
Computational LinguisticsLarge Language ModelsFact-checkingFake News
HY

Hongzhi Yin

Professor and ARC Future Fellow, University of Queensland
Recommender SystemGraph LearningSpatial-temporal PredictionEdge Intelligence
HY

Hwanjo Yu

POSTECH
data miningmachine learningrecommendation systemtime-series
XY

Xiaohu Yang

National University of Defense Technology
Plasma physicsLaser-plasma interactionInertial confinement fusionCharged particle beam
VG

Vivek Gupta

Assistant Professor of Computer Science, Arizona State University
Artificial IntelligenceNatural Language ProcessingLarge Language ModelsInformation Retrieval