Score
Designs and builds domain‑specific compilers and language toolchains, including domain‑specific languages, modeling formalisms, and architectures that translate high‑level domain abstractions into efficient executable representations. Analyzes and implements domain‑specific datasets, annotation and curation processes, evaluation metrics, pretraining and cross‑domain transfer strategies, data and model integration, and cross‑domain troubleshooting, collaboration, and communication to validate, deploy, and maintain the resulting toolchain.
Traditional compilers face limitations in development accessibility, optimization capabilities, and application scope. This work proposes the first multidimensional classification framework for large language model (LLM)-driven compilation, offering a systematic survey of existing research through four analytical dimensions: design philosophy, methodology, level of code abstraction, and task type. The study identifies three core design paradigms—Selector, Translator, and Generator—and highlights three transformative directions: democratizing compiler development, discovering novel optimization strategies, and expanding functional boundaries. It further argues that hybrid systems represent a critical pathway forward and provides a technical roadmap for building correct, scalable, and intelligent compilation tools.
This study investigates whether domain-specific languages (DSLs) enhance developers’ comprehension of data pipeline program structure. Method: A mixed-methods approach is employed—controlled experiments measure task accuracy, while structured surveys and qualitative coding analyze DSLs’ impact on domain experts’ structural awareness, accessibility, and alignment with mental models. Contribution/Results: This work provides the first empirical validation of systematic improvements in structural understanding of data pipelines afforded by DSLs. Results show statistically significant gains in comprehension accuracy (p < 0.01), driven by DSLs’ capacity to reinforce global program overviews, enforce syntactically constrained structures, and better align with users’ domain-specific mental models. Furthermore, DSLs lower the barrier to entry for programmers with limited experience, facilitate cross-tool knowledge transfer, and strengthen perception of dataflow structure.
The lack of systematic understanding of domain-specific language (DSL) evolution hinders the advancement of model-driven engineering (MDE) methods and tooling. Method: This study conducts the first large-scale empirical analysis of 1,002 Xtext-based textual DSL projects on GitHub, identifying 226 mature DSLs spanning 18 application domains. We propose a hybrid methodology integrating GitHub API mining, manual classification, and DSL metamodel analysis to quantify grammar coverage and characterize change types. Contribution/Results: We find that DSLs in domains such as Data Management exhibit broad adoption, rapid evolution, and long lifespans; grammar-driven development is the dominant paradigm, with Xtext frequently employed for refactoring existing languages. Among the 722 projects containing grammar definitions, only 33% provide textual examples, yet over 60% of grammar rules are empirically observed in use. Evolution is predominantly perfective—aimed at enhancing functionality and maintainability. The study delivers the first open-source DSL evolution dataset annotated with rich metadata, providing an empirical foundation for DSL engineering practice and tool development.
To address the opacity of code semantics in AI-assisted programming—hindering visual inspection and formal verification—this paper proposes a DSL-driven multimodal interaction framework. It anchors program semantics in a domain-specific language (e.g., Lingua Franca), integrates natural language and speech input, and constructs interpretable, visual program models. Real-time graphical rendering and staged refinement enable dynamic traceability throughout code generation. Model checking is embedded to ensure semantic consistency via formal verification. Implemented as a VS Code extension prototype, the framework maintains high code-generation quality while significantly enhancing developers’ understanding of and trust in AI behavior. The core contribution lies in the deep synergy among DSL-based modeling, multimodal interaction, and formal verification—achieving, for the first time in an IDE-integrated tool, a closed-loop workflow wherein AI-generated code is both semantically visualized and formally verifiable.
Modeling complex concurrent and timing-sensitive systems faces challenges in multi-objective compilation (for simulation, deployment, and formal verification), weak semantic consistency across targets, and the lack of expressive, unified modeling languages. Method: This paper introduces M, a textual modeling language grounded in the Actor model and discrete-event scheduling semantics, supporting temporal/state-triggered behaviors and asynchronous message passing. We design the first reusable, multi-target model compilation framework that uses M as a unified intermediate representation to enable semantics-preserving model transformations and code generation across heterogeneous targets. Contribution/Results: M serves as a common anchor for diverse domain-specific modeling languages, significantly enhancing model reusability and toolchain interoperability. The framework provides a general-purpose compilation infrastructure for heterogeneous system development—bridging simulation, implementation, and formal verification—while ensuring end-to-end semantic fidelity across compilation targets.
Large language models (LLMs) exhibit limited performance in domain-specific code generation—e.g., web, game, and mathematical programming—primarily due to insufficient semantic understanding of specialized APIs (e.g., React, Unity). This work presents the first systematic analysis revealing critical deficiencies in LLMs’ API-level cognition. To address this, we propose DomCoder, a domain-enhanced code generation framework that integrates three complementary API knowledge augmentation strategies: external knowledge retrieval, chain-of-thought (CoT) prompting, and CoT-aware fine-tuning. Evaluated across diverse domain-specific benchmarks, DomCoder achieves significant improvements in both functional correctness and domain-specific fidelity. Our results empirically validate that explicit API knowledge guidance effectively bridges the domain capability gap in LLMs, advancing their applicability to real-world software development tasks requiring deep platform expertise.
论文提出了一种领域导向的工具模式,通过选择预定义的、与领域对齐的参数化查询工具代替实时生成SQL,从而提高MCP服务器处理企业数据请求的效率和准确性。
This work addresses the lack of effective evaluation methods for assessing the scientific validity of domain-specific language (DSL) code—such as LAMMPS molecular dynamics input scripts—generated by large language models (LLMs). To tackle this challenge, the authors propose a lightweight validation framework that combines input file normalization, an extensible DSL parser, and static syntactic and semantic checks. This approach enables domain experts to efficiently verify LLM-generated outputs without requiring deep expertise in the target DSL. By circumventing costly runtime execution, the framework facilitates systematic benchmarking of mainstream LLMs on scientific DSL generation tasks, revealing their current limitations. The study thus provides a practical pathway toward the safe integration of LLMs into specialized scientific computing workflows.
This work addresses the challenge of domain modeling in privacy-sensitive industrial settings where closed-source large language models are inaccessible and locally deployed open-source small models are constrained by limited context windows, hindering direct extraction of domain knowledge from extensive codebases. To overcome this, the authors propose an iterative reasoning approach that integrates structural and semantic heuristics to prioritize and select critical code subsets, thereby guiding lightweight, locally deployable large language models to incrementally identify domain concepts and refine their boundaries—without requiring access to the full system context. This method represents the first integration of heuristic strategies with on-premise large language models for domain modeling, achieving high F1 scores across a benchmark dataset of ten real-world projects while preserving both privacy and model fidelity.
This work addresses the misalignment between domain models and code in Domain-Driven Design (DDD) caused by divergent evolution rhythms. To resolve this, the authors propose JDomInO, a bidirectional synchronization toolchain grounded in a shared metamodel that, for the first time, supports forward code generation and reverse refactoring across all twelve tactical DDD building blocks, thereby ensuring continuous consistency between Java implementations and tactical models. The approach integrates metamodel-driven round-trip engineering, Java static analysis, deterministic code generation, and model reconstruction techniques, while leveraging structured domain models as a precise contextual layer for AI-powered programming assistants. The forward path has been fully validated in a hotel management scenario, the reverse path logic verified through unit tests, and end-to-end validation is currently underway.
本文提出一种架构,通过分离意图解释、执行和解释,并基于领域本体约束分析链,解决了LLM辅助科学可视化中生成错误脚本的问题。