Score
Designs, builds, and operates systems and processes for collecting, storing, organizing, validating, securing, and distributing data across its lifecycle, including schemas, databases, data pipelines, metadata, cataloging, provenance, backup, and retention policies. Analyzes and enforces data quality, access control, integration, governance, and performance to ensure data is reliable, discoverable, consistent, and compliant for downstream use.
In highly regulated domains such as finance, data quality control (QC) is often fragmented into isolated preprocessing steps, undermining end-to-end trustworthy AI pipelines. To address this, we propose the first AI-driven DataOps framework that embeds QC as a system-level core component. Our framework deeply integrates rule-based engines, statistical analysis, and custom AI-powered anomaly detection across the entire data lifecycle—from ingestion and transformation to model deployment—enabling dynamic remediation, policy-configurable workflows, and end-to-end auditability. Technically, it unifies data profiling, stream processing, cloud-native storage interfaces, and a proprietary AI detection module. Evaluated in a real-world financial production environment, the framework achieves significantly improved anomaly recall, reduces manual intervention by 42%, and ensures audit completeness and full data traceability under high-throughput conditions, fully satisfying regulatory compliance requirements.
Ad hoc SQL development lacks engineering rigor, leading to data silos, logical redundancy, and ineffective data governance. Method: This paper proposes a DataOps-driven CI/CD framework for analytical SQL warehouses, featuring a novel five-stage automated pipeline—Lint, Optimize, Parse, Validate, Observe—that embeds quality assurance and enables end-to-end lifecycle governance. Contribution/Results: We introduce the DataOps Controls Scorecard and a requirements traceability matrix, explicitly mapping 12 governance criteria to CI/CD stages to ensure control completeness and scalability. The framework integrates Agile, Lean, and DevOps principles with static analysis, syntactic parsing, optimization recommendations, validation testing, and observability. Empirical evaluation demonstrates significant improvements in data quality, development transparency, and cross-functional collaboration, providing a sustainable, production-ready pathway for large-scale analytical systems.
To address insufficient semantic description of multidimensional aggregate/summary data, poor adaptability of metadata standards, and cross-source interoperability challenges in big data environments, this paper proposes a multidimensional data source profiling metadata model tailored for data ecosystems. Built upon RDF, the model is the first to support extensible semantic modeling of both aggregate and summary multidimensional data, enabling semantic alignment of dimensions and measures with reference knowledge graphs. It integrates multi-granularity metadata profiles—spanning source-level, attribute-level, and value-distribution characteristics. The model ensures flexible extensibility and cross-source interoperability. Experimental results demonstrate that profile generation time scales linearly with data cardinality, confirming its engineering practicality and predictable performance.
This study addresses the strategic lock-in and operational risks organizations face when relying on commercial data intermediaries to ensure the timeliness and reliability of master data. It pioneers the systematic integration of Self-Sovereign Identity (SSI) into master data management by synthesizing insights from hermeneutic literature review, expert interviews, and design science research methodologies. The resulting design theory embeds a trustworthy master data management framework within a data space reference architecture, emphasizing data sovereignty, reliability, and accountability. Validated through evaluation by industry experts, the proposed framework enables trusted, controllable data sharing and governance within data ecosystems.
This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.
Existing data pipelines often suffer from weak governance, leading to delayed schema validation, inconsistent cross-language execution, and misalignment with business semantics. This work proposes treating data contracts as types, leveraging the “everything-as-code” paradigm to inject schema annotations—encompassing column types, constraints, documentation, and lineage—into input and output tables within a lakehouse architecture via multi-language SDKs. These annotations are parsed across multiple phases of the execution lifecycle, deeply integrating data contracts into the type system. The approach enables both deterministic and non-deterministic reasoning over data flows across languages and execution engines, significantly enhancing the reliability of production data pipelines and ensuring consistent interoperability across systems.
In the era of large language models, traditional record-centric data engineering struggles to meet the demand for organizational knowledge as executable infrastructure. This work proposes a novel paradigm—knowledge architecture—that systematically reimagines core data engineering mechanisms by upgrading ETL, data lineage, and catalogs into knowledge ingestion, change detection, provenance, and knowledge catalogs. It introduces knowledge views and a three-tier layered model (raw–refined–operational) to structure knowledge effectively. By integrating emerging standards such as LLM Wiki and Open Knowledge Format (OKF), this study formally defines knowledge architecture for the first time and establishes a theoretical framework that supports knowledge representation, governance, and operational delivery, enabling direct invocation of organizational knowledge by humans, agents, workflows, and models alike.
This work addresses the inadequacy of existing large language model (LLM) lifecycle frameworks, which predominantly emphasize operational efficiency while lacking explicit support for security-critical activities—such as data provenance, component signing, and access control—and failing to align governance requirements with specific lifecycle phases. The paper proposes the first security-oriented LLM system lifecycle model, structured not by workflow but by security boundaries, organizing 32 phases into four layered pipelines: data, model, distribution, and application, while integrating LLMOps and governance pillars. It uniquely identifies 13 distinct security-critical phases and exposes a structural imbalance wherein regulatory evidence is concentrated at deployment despite pivotal decisions occurring during development. By mapping key standards—including NIST AI RMF, the EU AI Act, and ISO/IEC 42001—the study establishes a phase-to-governance correspondence mechanism, yielding a comprehensive, lifecycle-spanning security analysis framework that offers structured guidance for compliance and secure design.