Score
Designs and builds processes, pipelines, and datasets to discover and define a target domain, collect and curate representative domain data, and analyze domain characteristics to inform downstream models or studies. Tasks include specifying data sources and collection protocols, establishing annotation and quality‑control and curation workflows, and producing documentation and curated datasets that capture the domain's scope and constraints.
This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.
This study investigates whether domain-specific languages (DSLs) enhance developers’ comprehension of data pipeline program structure. Method: A mixed-methods approach is employed—controlled experiments measure task accuracy, while structured surveys and qualitative coding analyze DSLs’ impact on domain experts’ structural awareness, accessibility, and alignment with mental models. Contribution/Results: This work provides the first empirical validation of systematic improvements in structural understanding of data pipelines afforded by DSLs. Results show statistically significant gains in comprehension accuracy (p < 0.01), driven by DSLs’ capacity to reinforce global program overviews, enforce syntactically constrained structures, and better align with users’ domain-specific mental models. Furthermore, DSLs lower the barrier to entry for programmers with limited experience, facilitate cross-tool knowledge transfer, and strengthen perception of dataflow structure.
Event logs frequently contain noise and missing entries, while conventional process discovery methods neglect domain knowledge, resulting in biased models and low downstream reliability. To address this, we propose the first interactive process discovery framework integrating large language models (LLMs). Our approach employs prompt engineering to elicit declarative process rules from expert-provided natural language descriptions; these rules are jointly processed with event logs by an enhanced Inductive Miner revised (IMr) algorithm to recursively construct process models that balance accuracy and interpretability. The system enables real-time expert feedback and iterative rule refinement. Empirical evaluation demonstrates substantial improvements in model adaptability and structural soundness. Expert assessments confirm high usability and practical deployability, validating the framework’s effectiveness in bridging domain expertise with automated process discovery.
How can domain experts’ tacit knowledge—regarding data provenance, quality, and usage—be efficiently elicited to improve domain adaptability in visualization design? This paper introduces the “Data Therapist” paradigm: an LLM-driven web tool integrating hybrid active questioning and interactive annotation. It supports multi-granularity structured annotation and iterative follow-up queries to systematically externalize and model tacit knowledge. The method synergizes large language models, interactive knowledge elicitation interfaces, and qualitative user studies. Empirical validation across molecular biology, accounting, political science, and usable security reveals cross-domain patterns in data reasoning. Results demonstrate significant improvements in visualization systems’ understanding and support of domain semantics, enabling more robust, domain-informed visualization design. By formalizing and structuring expert knowledge, this work establishes a scalable, reusable knowledge infrastructure for data-driven, automated visualization generation.
Public datasets in the LLM4RE (Large Language Models for Requirements Engineering) domain are fragmented, poorly documented, and lack systematic description, hindering comparability and reuse. Method: We conduct the first systematic dataset mapping study in LLM4RE, analyzing 62 publicly available datasets drawn from 43 scholarly publications along dimensions including document type, granularity, RE task phase, domain, and language. We propose the first domain-specific dataset classification and characterization framework for LLM4RE. Contribution/Results: Our framework identifies critical research gaps—particularly in requirements elicitation, requirements management, and multilingual support—and we release an open dataset catalog alongside a standardized featureization schema. This work significantly enhances dataset visibility, structural consistency, and cross-study comparability, laying the foundation for a unified benchmarking repository in LLM4RE.
本文提出一种架构,通过分离意图解释、执行和解释,并基于领域本体约束分析链,解决了LLM辅助科学可视化中生成错误脚本的问题。
This study addresses the challenge that domain experts, due to limited query language proficiency, often struggle to independently conduct context-specific data quality analyses and must rely on technical specialists, resulting in inefficient workflows. To overcome this limitation, the paper proposes the Quality Pattern Model (QPM) framework—a novel, template-based mechanism that is agnostic to both database technologies and application domains, enabling non-technical users to autonomously define data quality analysis logic. Leveraging a model-driven approach, the authors implement QPM prototypes over XML, RDF, and Neo4j. Experimental results demonstrate that QPM’s expressiveness matches or exceeds that of mainstream query languages while significantly enhancing domain experts’ analytical autonomy. The framework’s effectiveness has been validated in the cultural heritage domain.
This study addresses the lack of empirical analysis on the large-scale adoption of Domain-Driven Design (DDD) in real-world open-source projects, where its prevalence, architectural patterns, and technology preferences remain unclear. By mining GitHub repositories and applying an initial screening based on topic tags and README keywords, we introduce a semantic validation pipeline powered by GPT-4o along with a triple-majority voting mechanism to construct the first large-scale dataset of verified DDD projects (2,502 in total). Our analysis reveals that DDD adoption has accelerated since 2017, with DDD-based projects exhibiting significantly longer lifespans than average. Layered and Clean Architectures dominate, while CQRS and event sourcing are primarily employed in distributed systems. Notably, C# and TypeScript lead in usage, challenging the assumption of Java’s centrality, and 25.3% of projects lack explicitly defined bounded contexts.
研究解决数据仓库中缺乏变量级元数据的问题,通过提出一种与DDI-CDI模型兼容的元数据应用配置文件方法来丰富元数据并进行模型-数据一致性检查。
Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.