domain discovery

Designs and builds processes, pipelines, and datasets to discover and define a target domain, collect and curate representative domain data, and analyze domain characteristics to inform downstream models or studies. Tasks include specifying data sources and collection protocols, establishing annotation and quality‑control and curation workflows, and producing documentation and curated datasets that capture the domain's scope and constraints.

domaindiscovery

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.46
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

Can a domain-specific language improve program structure comprehension of data pipelines? A mixed-methods study

May 22, 2025
PH
Philip Heltweg
🏛️ Friedrich-Alexander-Universität Erlangen-Nürnberg

This study investigates whether domain-specific languages (DSLs) enhance developers’ comprehension of data pipeline program structure. Method: A mixed-methods approach is employed—controlled experiments measure task accuracy, while structured surveys and qualitative coding analyze DSLs’ impact on domain experts’ structural awareness, accessibility, and alignment with mental models. Contribution/Results: This work provides the first empirical validation of systematic improvements in structural understanding of data pipelines afforded by DSLs. Results show statistically significant gains in comprehension accuracy (p < 0.01), driven by DSLs’ capacity to reinforce global program overviews, enforce syntactically constrained structures, and better align with users’ domain-specific mental models. Furthermore, DSLs lower the barrier to entry for programmers with limited experience, facilitate cross-tool knowledge transfer, and strengthen perception of dataflow structure.

Does a domain-specific language enhance data pipeline program structure comprehension?How does a domain-specific language affect correctness and efficiency in understanding pipelines?What human factors influence domain experts' use of domain-specific languages for pipelines?

Event logs frequently contain noise and missing entries, while conventional process discovery methods neglect domain knowledge, resulting in biased models and low downstream reliability. To address this, we propose the first interactive process discovery framework integrating large language models (LLMs). Our approach employs prompt engineering to elicit declarative process rules from expert-provided natural language descriptions; these rules are jointly processed with event logs by an enhanced Inductive Miner revised (IMr) algorithm to recursively construct process models that balance accuracy and interpretability. The system enables real-time expert feedback and iterative rule refinement. Empirical evaluation demonstrates substantial improvements in model adaptability and structural soundness. Expert assessments confirm high usability and practical deployability, validating the framework’s effectiveness in bridging domain expertise with automated process discovery.

Current methods lack integration of natural language domain constraintsEvent logs alone produce unreliable models due to incompleteness and noiseProcess discovery models often ignore valuable domain knowledge from experts

Data Therapist: Eliciting Domain Knowledge from Subject Matter Experts Using Large Language Models

May 01, 2025
SS
Sungbok Shin
🏛️ Inria | Université Paris-Saclay | Seoul National University | Oregon State University | Aarhus University

How can domain experts’ tacit knowledge—regarding data provenance, quality, and usage—be efficiently elicited to improve domain adaptability in visualization design? This paper introduces the “Data Therapist” paradigm: an LLM-driven web tool integrating hybrid active questioning and interactive annotation. It supports multi-granularity structured annotation and iterative follow-up queries to systematically externalize and model tacit knowledge. The method synergizes large language models, interactive knowledge elicitation interfaces, and qualitative user studies. Empirical validation across molecular biology, accounting, political science, and usable security reveals cross-domain patterns in data reasoning. Results demonstrate significant improvements in visualization systems’ understanding and support of domain semantics, enabling more robust, domain-informed visualization design. By formalizing and structuring expert knowledge, this work establishes a scalable, reusable knowledge infrastructure for data-driven, automated visualization generation.

Capturing tacit data context through interactive annotationEliciting domain knowledge from experts for visualizationImproving visualization design with AI-supported expert insights

ShaRE your Data! Characterizing Datasets for LLM-based Requirements Engineering

Oct 21, 2025
QM
Quim Motger
🏛️ Universitat Politècnica de Catalunya

Public datasets in the LLM4RE (Large Language Models for Requirements Engineering) domain are fragmented, poorly documented, and lack systematic description, hindering comparability and reuse. Method: We conduct the first systematic dataset mapping study in LLM4RE, analyzing 62 publicly available datasets drawn from 43 scholarly publications along dimensions including document type, granularity, RE task phase, domain, and language. We propose the first domain-specific dataset classification and characterization framework for LLM4RE. Contribution/Results: Our framework identifies critical research gaps—particularly in requirements elicitation, requirements management, and multilingual support—and we release an open dataset catalog alongside a standardized featureization schema. This work significantly enhances dataset visibility, structural consistency, and cross-study comparability, laying the foundation for a unified benchmarking repository in LLM4RE.

Addressing limited visibility of datasets in LLM4RE researchCharacterizing fragmented datasets for LLM-based Requirements EngineeringInvestigating under-represented RE tasks and dataset descriptors

Latest Papers

What's happening recently
View more

This study addresses the challenge that domain experts, due to limited query language proficiency, often struggle to independently conduct context-specific data quality analyses and must rely on technical specialists, resulting in inefficient workflows. To overcome this limitation, the paper proposes the Quality Pattern Model (QPM) framework—a novel, template-based mechanism that is agnostic to both database technologies and application domains, enabling non-technical users to autonomously define data quality analysis logic. Leveraging a model-driven approach, the authors implement QPM prototypes over XML, RDF, and Neo4j. Experimental results demonstrate that QPM’s expressiveness matches or exceeds that of mainstream query languages while significantly enhancing domain experts’ analytical autonomy. The framework’s effectiveness has been validated in the cultural heritage domain.

data qualitydomain expertsquality analysis

This study addresses the lack of empirical analysis on the large-scale adoption of Domain-Driven Design (DDD) in real-world open-source projects, where its prevalence, architectural patterns, and technology preferences remain unclear. By mining GitHub repositories and applying an initial screening based on topic tags and README keywords, we introduce a semantic validation pipeline powered by GPT-4o along with a triple-majority voting mechanism to construct the first large-scale dataset of verified DDD projects (2,502 in total). Our analysis reveals that DDD adoption has accelerated since 2017, with DDD-based projects exhibiting significantly longer lifespans than average. Layered and Clean Architectures dominate, while CQRS and event sourcing are primarily employed in distributed systems. Notably, C# and TypeScript lead in usage, challenging the assumption of Java’s centrality, and 25.3% of projects lack explicitly defined bounded contexts.

architectural practicesbusiness contextDomain-Driven Design

Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.

heterogeneous datareproducibilityschema

Hot Scholars

XC

Xueqi Cheng

Ph.D. student, Florida State University
Data miningLLMGNNComputational social science
SA

Sophia Ananiadou

Professor, Computer Science, Manchester University, National Centre for Text Mining
Natural Language ProcessingText MiningComputational LinguisticsArtificial Intelligence
MK

Marcos Kalinowski

Professor, Pontifical Catholic University of Rio de Janeiro (PUC-Rio)
Empirical Software EngineeringAI EngineeringAI4SEHuman Aspects in Software Engineering
QL

Qinghua Lu

Group Leader, Software Systems Research Group, CSIRO's Data61
AI EngineeringSE4AISoftware ArchitectureAI Safety
LW

Lijun Wu

Shanghai AI Laboratory
MLLLMAI4Science