data governance

Designs, implements, and evaluates the policies, roles, processes, standards, and technical controls that govern the availability, quality, integrity, security, and lifecycle management of an organization’s data assets. This work includes defining ownership and stewardship, metadata and cataloging practices, access and sharing rules, data quality and lineage measures, and compliance and risk controls.

datagovernance

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.49
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the lack of a unified standard for assessing enterprise data asset quality and the unclear mechanisms linking such quality to value realization. It proposes an integrative three-dimensional framework—encompassing management capability, standards compliance, and benefit realization—and uniquely combines grounded theory, LDA topic modeling, PLS-SEM, necessary condition analysis (NCA), and fuzzy-set qualitative comparative analysis (fsQCA) to simultaneously examine structural relationships, necessary conditions, and configurational pathways. The findings confirm significant positive relationships among the three dimensions, identify necessary conditions for high-quality data assets, and uncover multiple equifinal configurations that reveal a dual-path chain mechanism driven by both governance orientation and benefit-driven imperatives.

Data Asset QualityData Quality EvaluationEnterprise Data Governance

DataOps-driven CI/CD for analytics repositories

Nov 15, 2025
DV
Dmytro Valiaiev
🏛️ University of Arkansas Little Rock

Ad hoc SQL development lacks engineering rigor, leading to data silos, logical redundancy, and ineffective data governance. Method: This paper proposes a DataOps-driven CI/CD framework for analytical SQL warehouses, featuring a novel five-stage automated pipeline—Lint, Optimize, Parse, Validate, Observe—that embeds quality assurance and enables end-to-end lifecycle governance. Contribution/Results: We introduce the DataOps Controls Scorecard and a requirements traceability matrix, explicitly mapping 12 governance criteria to CI/CD stages to ensure control completeness and scalability. The framework integrates Agile, Lean, and DevOps principles with static analysis, syntactic parsing, optimization recommendations, validation testing, and observability. Empirical evaluation demonstrates significant improvements in data quality, development transparency, and cross-functional collaboration, providing a sustainable, production-ready pathway for large-scale analytical systems.

Addressing ad-hoc SQL development lacking software engineering rigorProviding standardized DataOps framework for analytics pipeline managementSolving data governance challenges and validation impossibility in analytics

Large organizations struggle to sustain information security and regulatory compliance in dynamic, evolving environments. Method: This study models enterprise information security governance as a multidimensional dynamical system and, for the first time, formalizes it as a feedback regulation problem within control-theoretic frameworks. Leveraging the UK BS standard, we construct an enterprise-scale digital twin with 1.2 million parameters and propose a quantification paradigm centered on an integral-type security state metric, enabling real-time security态势 characterization and closed-loop compliance verification. Contribution/Results: The work transcends traditional static audit paradigms by establishing a novel digital twin–enabled security governance approach—standards-driven, parameter-auditable, quantitatively evaluable, and response-controllable. The solution has been fully deployed across an operational enterprise and integrated with organization-wide capability alignment, yielding significant improvements in security resilience and regulatory response efficiency.

Access ControlInformation SecuritySecurity Measures Evaluation

Existing Web data storage platforms struggle to meet the demands of decentralized, semantically rich, and legally compliant data usage control. This work proposes a novel approach that integrates the User-Managed Access (UMA) authorization framework with the W3C Open Digital Rights Language (ODRL) policy language to replace Solid’s native access control mechanism, thereby decoupling authorization from storage. For the first time within the Solid ecosystem, this integration advances access control from mere permission management toward legally aware usage control. The authors also design a policy evaluation mechanism tailored for non-standardized semantic environments. A prototype implementation demonstrates that the proposed method maintains compatibility with Solid while enabling flexible, interoperable, and legally aligned data governance.

access controldata governancedecentralized data ecosystems

Latest Papers

What's happening recently
View more

This work addresses the inadequacy of existing large language model (LLM) lifecycle frameworks, which predominantly emphasize operational efficiency while lacking explicit support for security-critical activities—such as data provenance, component signing, and access control—and failing to align governance requirements with specific lifecycle phases. The paper proposes the first security-oriented LLM system lifecycle model, structured not by workflow but by security boundaries, organizing 32 phases into four layered pipelines: data, model, distribution, and application, while integrating LLMOps and governance pillars. It uniquely identifies 13 distinct security-critical phases and exposes a structural imbalance wherein regulatory evidence is concentrated at deployment despite pivotal decisions occurring during development. By mapping key standards—including NIST AI RMF, the EU AI Act, and ISO/IEC 42001—the study establishes a phase-to-governance correspondence mechanism, yielding a comprehensive, lifecycle-spanning security analysis framework that offers structured guidance for compliance and secure design.

governance frameworklarge language modelsLLM systems

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

Hot Scholars

JH

J. Harry Caufield

Lawrence Berkeley National Laboratory
knowledge graphsbiomedical informaticsartificial intelligencelarge language models
PN

Preslav Nakov

Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
Computational LinguisticsLarge Language ModelsFact-checkingFake News
QL

Quanzheng Li

Massachusetts General Hospital, Harvard Medical School
Image ReconstructionMedical Image AnalysisDeep Learning in MedicineMultimodality Medical Data Analysis
RM

Rashid Mushkani

University of Montreal I Mila
Public (Space & Life)Sociotechnical AIUrban AnalyticsCommunity-Centered AI
WN

Wei Ni

FIEEE, AAIA Fellow, Senior Principal Scientist & Conjoint Professor, CSIRO/UNSW
6G security and privacyconnected and trusted intelligenceapplied AI/ML