Score
Design domain-specific languages: create and specify programming or configuration languages tailored to a particular problem domain by defining concrete and abstract syntax, formal grammars, and the language's operational or denotational semantics. Build the accompanying artifacts and tools — parsers, analyzers, compilers/interpreters, code generators or transformers — and iterate the grammar and syntax to validate expressiveness, readability, modularity, and correct mapping or compatibility with target implementations or platforms.
This study investigates whether domain-specific languages (DSLs) enhance developers’ comprehension of data pipeline program structure. Method: A mixed-methods approach is employed—controlled experiments measure task accuracy, while structured surveys and qualitative coding analyze DSLs’ impact on domain experts’ structural awareness, accessibility, and alignment with mental models. Contribution/Results: This work provides the first empirical validation of systematic improvements in structural understanding of data pipelines afforded by DSLs. Results show statistically significant gains in comprehension accuracy (p < 0.01), driven by DSLs’ capacity to reinforce global program overviews, enforce syntactically constrained structures, and better align with users’ domain-specific mental models. Furthermore, DSLs lower the barrier to entry for programmers with limited experience, facilitate cross-tool knowledge transfer, and strengthen perception of dataflow structure.
This work addresses the significant performance degradation of large language models (LLMs) in generating code for constraint-based domain-specific languages (DSLs), such as OCL and Alloy, and the absence of systematic evaluation methodologies. The paper introduces the first evaluation framework tailored for constraint DSL code generation, which systematically assesses LLM capabilities in translating natural language to DSL through both syntactic correctness and semantic accuracy, leveraging formal verification. Experimental comparisons across Python, OCL, and Alloy reveal that LLMs perform markedly better on general-purpose languages, that models with limited context windows struggle to jointly generate constraints and domain models, and that incorporating code repair and multi-candidate generation strategies substantially improves output quality. The framework further enables systematic analysis of prompting templates, repair mechanisms, and multi-turn generation strategies.
Formal program specifications are notoriously difficult, error-prone, and inefficient to write manually. To address this, we propose a two-stage LLM-driven approach: dialogue-guided specification synthesis followed by mutation-based verification. First, multi-turn dialogues model complex semantic requirements; second, four mutation operators—insertion, replacement, deletion, and reordering—enable verifiability-driven selection, eliminating reliance on rigid templates or syntactic grammars. Our method integrates code understanding, prompt engineering, and heuristic verifiability assessment. Evaluated on SV-COMP and a custom Java benchmark comprising 385 programs, it generates 279 verifiable specifications. These achieve significantly higher completeness and accuracy than pure-LLM baselines and classical tools (e.g., Houdini, Daikon). To our knowledge, this is the first approach to achieve both high coverage and formal verifiability in fully automated specification generation.
Large language models (LLMs) exhibit limited performance in domain-specific code generation—e.g., web, game, and mathematical programming—primarily due to insufficient semantic understanding of specialized APIs (e.g., React, Unity). This work presents the first systematic analysis revealing critical deficiencies in LLMs’ API-level cognition. To address this, we propose DomCoder, a domain-enhanced code generation framework that integrates three complementary API knowledge augmentation strategies: external knowledge retrieval, chain-of-thought (CoT) prompting, and CoT-aware fine-tuning. Evaluated across diverse domain-specific benchmarks, DomCoder achieves significant improvements in both functional correctness and domain-specific fidelity. Our results empirically validate that explicit API knowledge guidance effectively bridges the domain capability gap in LLMs, advancing their applicability to real-world software development tasks requiring deep platform expertise.
This work addresses the problem of program synthesis in domain-specific languages (DSLs) that involve numeric constants and require optimization of quantitative objectives such as accuracy. The authors propose a provably optimal search method that constructs a search graph over program subsets and integrates A* search with a heuristic derived from abstract interpretation to efficiently prune suboptimal subtrees. The key innovation lies in the design of abstract transformers tailored to DSL components with monotonic semantics, enabling a pruning mechanism that guarantees optimality. Experimental evaluation on two real-world DSLs demonstrates that the approach substantially outperforms existing state-of-the-art synthesizers, achieving significant improvements in scalability while maintaining correctness and optimality guarantees.
This work addresses the challenge that general-purpose large language models often struggle to accurately invoke library functions and adhere to domain-specific conventions when generating code for specialized frameworks such as Scikit-learn and OpenCV. To systematically evaluate customization strategies, the authors construct a synthetic programming dataset spanning general Python, Scikit-learn, and OpenCV, and assess three approaches—few-shot prompting, retrieval-augmented generation (RAG), and LoRA fine-tuning—within a unified framework. The study presents the first comparative analysis of prompt engineering versus parameter-efficient fine-tuning in domain-specific code generation, revealing practical trade-offs among accuracy, cost, and flexibility. Experimental results demonstrate that LoRA fine-tuning significantly outperforms prompt-based methods in both accuracy and domain alignment, whereas few-shot prompting and RAG, while improving relevance, offer limited gains in overall correctness.
Manually configuring linters requires expert knowledge and struggles to adapt across multiple programming languages, coding standards, and tooling ecosystems, leading to high maintenance overhead. This work proposes LintCFG, the first approach to apply compiler design principles to automated linter configuration generation. It introduces a tool-agnostic domain-specific language (DSL) to structurally encode coding rules and leverages large language models to automatically compile natural language specifications into concrete linter configurations, enabling end-to-end automation across languages, standards, and tools. Evaluated on Java Checkstyle tasks, the DSL achieves over 90% precision and recall in rule representation, with fine-grained configuration generation exceeding 70% accuracy—more than doubling the performance of baseline methods. User studies confirm significant gains in developer productivity, and the approach successfully generalizes to JavaScript ESLint scenarios.
This work addresses the challenge of deploying large language models in production settings, where they often fail to meet low-latency requirements, while smaller models typically suffer from limited reasoning capabilities, hallucinations, and insufficient long-context memory. To overcome these limitations, the authors propose supervised fine-tuning small models such as Mistral on domain-specific natural language–code paired data, thereby internalizing domain knowledge directly into model weights and substantially reducing reliance on runtime context. Experimental results demonstrate that the fine-tuned small models outperform larger counterparts in code generation quality while maintaining lower latency. Load testing and real-world deployment confirm their efficiency and stability. Furthermore, the approach supports additional customer-specific fine-tuning without compromising general-purpose capabilities, offering a practical pathway toward efficient and accurate domain-specific code generation.
This work proposes a structured prompting framework leveraging large language models (LLMs) to reduce the human effort and complexity inherent in Domain-Driven Design (DDD) implementation. The approach decomposes the DDD process into five sequential steps: event storming simulation, bounded context identification, aggregate design, glossary generation, and technical architecture mapping. It represents the first systematic application of prompt engineering across the entire DDD workflow, positioning the LLM as an expert collaborator rather than a replacement for human designers. Experimental results demonstrate that the first three steps effectively produce high-quality design artifacts—such as domain glossaries and context maps—whereas the latter two steps suffer from error propagation, thereby underscoring the necessity of human-in-the-loop collaboration for critical design decisions.
This work presents the first systematic investigation into the capability of large language models (LLMs) to generate program specifications involving higher-order logical constructs, which are essential for expressing complex verification properties yet remain beyond the reach of existing LLMs that predominantly handle basic syntactic forms. The authors design four syntactic configurations spanning different levels of abstraction and establish a comprehensive evaluation framework to assess a range of representative LLMs on standard verification benchmarks. Experimental results demonstrate that LLMs can effectively produce valid higher-order logical expressions; moreover, integrating logical constructs with base syntax significantly enhances verification efficacy and robustness without substantially increasing verification overhead. The study also reveals distinct advantages of two refinement paradigms in specification generation.