define data requirements

Designs and documents the specific data elements, attributes, types, formats, allowable values, semantics, quality requirements, and validation rules required to satisfy stakeholder needs. Breaks down high-level requirements into concrete field-level specifications and formal data standards—naming conventions, schemas, metadata and exchange formats—to enable consistent collection, storage, validation and interoperability.

definedatarequirements

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.76
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of operationalizing GDPR compliance in software engineering—specifically, how to realize “Privacy by Design” (PbD) at the requirements and system specification levels while reconciling heterogeneous stakeholder interests and ensuring semantic consistency and traceability between legal provisions and technical specifications. We propose a formal modeling approach grounded in original legal concepts, systematically mapping GDPR articles to reusable privacy requirement patterns. Integrating systematic literature analysis, industry interviews, and requirements modeling, we develop a joint specification framework supporting cross-layer abstraction and transparent, bidirectional traceability. Empirical evaluation demonstrates that the framework significantly improves the accuracy of privacy requirement elicitation and the transparency of regulatory specification, thereby providing a scalable, methodology-driven foundation for law–technology co-governance.

Aligning GDPR requirements with software engineering specificationsBridging problem-solution space gap in privacy by designCapturing legal knowledge in system specifications for compliance

Machine-interpretable Engineering Design Standards for Valve Specification

Oct 02, 2025
AG
Anders Gjerver
🏛️ Aibel AS | University of Western Australia | Equinor | DNV

Engineering design standards—typically expressed in natural language and tabular formats—are difficult for machines to interpret and validate automatically. Method: This paper proposes a modular ontology modeling approach grounded in the ISO/IEC/IEEE 24765 (IDO) top-level ontology, transforming textual and tabular specifications from standards such as ISO into OWL-based, W3C-compliant executable semantic ontologies, and integrating them with the ISO DIS 23726-3 Industrial Data Ontology. The resulting ontologies enable semantic reasoning and automated design rule verification. Results: The method achieves, for the first time, automated compliance checking against international materials and piping standards—including ASME B16.34 and ISO 15761—during valve selection. Its core contribution is a reusable, extensible semantic asset model that closes the loop from standard documents → machine-interpretable ontologies → design quality assurance, providing a practical, scalable pathway for standards development organizations to advance toward digital and intelligent transformation.

Automating quality assurance for plant design and equipment selectionEnabling semantic validation of valve specifications against industry standardsTransforming engineering design standards into machine-interpretable ontologies

This study addresses the challenges posed by divergent and conflicting data protection regulations across jurisdictions, which hinder the early identification of compliance requirements in software development and often lead to costly rework and legal risks. Drawing on interviews with 70 legal experts from G20 and other countries, the research employs systematic content analysis and deductive qualitative methods to distill, for the first time from a legal expert perspective, both commonalities—such as consent—and key divergences—such as the right to be forgotten—across global data protection laws. These insights are innovatively operationalized into a comprehensive set of Data Protection Officer (DPO) user stories mapped to each phase of the software development lifecycle and enterprise architecture layers, significantly enhancing the actionable integration of compliance requirements into early-stage software engineering practices.

data protection regulationsprivacy complianceregulatory data protection requirements

In software design, paradigm-implied semantic expectations—such as data abstraction consistency and feedback-control closed-loop behavior—are often left implicit, leading to design deviations and verification challenges. To address this, we introduce the concept of *design obligations*: explicit, logically formalizable, and verifiable specifications that codify such implicit constraints inherent to design paradigms. Leveraging formal modeling and paradigm semantics analysis, we establish two obligation frameworks—one for data-abstraction-based systems and another for feedback-driven adaptive systems—precisely capturing their core semantic requirements. We demonstrate that common design flaws stem from obligation violations and show how these obligations enable rigorous compliance verification and pedagogical application. This work bridges the semantic gap between design intent and implementation, providing both theoretical foundations and a methodological framework for paradigm-driven design assurance.

Addressing implicit or informal design expectations in software paradigms.Ensuring software designs meet semantic expectations beyond syntax.Introducing 'design obligations' to enforce proper paradigm use.

Synthesizing JSON Schema Transformers

May 27, 2024
JS
Jack Stanek
🏛️ University of Wisconsin - Madison

To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.

Automating transformation between different JSON Schema versionsGenerating programs to convert JSON data between schemasPreventing data loss during JSON Schema evolution

Latest Papers

What's happening recently
View more

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

This study addresses the lack of systematic methodologies for selecting data architectures in modern organizations grappling with vast, heterogeneous data environments. To this end, it proposes the DATER conceptual framework, which establishes a unified taxonomy of technical requirements and systematically examines the historical evolution, core characteristics, and applicability boundaries of six prominent data architectures: data warehouses, data lakes, lakehouses, data fabrics, and data meshes. Through conceptual modeling and multidimensional comparative analysis, the framework clarifies overlaps and distinctions among these architectures, articulating their respective strengths and limitations. By offering a structured evaluation tool, DATER significantly enhances the strategic alignment and contextual appropriateness of data architecture design for both researchers and practitioners.

data architecturedata integrationdata management

This study addresses the challenge in attributed graph schema design of whether repeatedly occurring descriptive attributes should be embedded within nodes or externalized as reusable metadata. Building upon Fifth Normal Form (5NF), the authors propose a principled decision framework that systematically identifies metadata candidates based on semantic criteria rather than mere repetition frequency. The approach classifies attributes into characteristic nodes, embedded properties, or borderline cases using five key principles: cross-element occurrence frequency, conceptual independence, lossless externalizability, reuse potential, and governance relevance. Empirical validation through a library domain case study and an entity classification task demonstrates that repetition alone is insufficient for externalization decisions—semantic judgment is essential. The proposed method significantly enhances the accuracy, consistency, and reusability of metadata modeling in graph-based systems.

embedded propertiesmetadataproperty graph schemas

This study addresses the challenge that domain experts, due to limited query language proficiency, often struggle to independently conduct context-specific data quality analyses and must rely on technical specialists, resulting in inefficient workflows. To overcome this limitation, the paper proposes the Quality Pattern Model (QPM) framework—a novel, template-based mechanism that is agnostic to both database technologies and application domains, enabling non-technical users to autonomously define data quality analysis logic. Leveraging a model-driven approach, the authors implement QPM prototypes over XML, RDF, and Neo4j. Experimental results demonstrate that QPM’s expressiveness matches or exceeds that of mainstream query languages while significantly enhancing domain experts’ analytical autonomy. The framework’s effectiveness has been validated in the cultural heritage domain.

data qualitydomain expertsquality analysis

This study addresses the challenge of transforming stakeholder requirements into product requirements in software-driven automotive systems. Leveraging a dataset of 8,082 stakeholder requirements and 5,870 product requirements provided by Infineon, the research employs a hybrid methodology integrating structural statistics, decision modeling, traceability mining, textual analysis, and hardware-software linkage to systematically analyze the requirement refinement process. It reveals, for the first time, that requirement complexity primarily stems from ambiguous architectural scope and missing contextual information rather than linguistic redundancy. The work establishes a classification framework for mapping stakeholder to product requirements, identifies systematic differences across abstraction levels, and proposes key improvements in requirement validation, deviation management, and contextual tooling to support efficient and reusable automotive development.

automotive industryproduct requirementsrequirement engineering

Hot Scholars

PL

Patricia Lago

Full Professor, S2 Group, Dept. Computer Science, Vrije Universiteit Amsterdam
Software ArchitectureSoftware EngineeringSoftware SustainabilityGreen Software
BI

Boris Ivanovic

NVIDIA
Machine LearningDeep LearningComputer VisionRobotics