data governance

Designing and operationalizing policies, processes, and controls for collecting, sharing, preprocessing, and maintaining data so it remains high-quality, ethically managed, and legally compliant (e.g., GDPR). This includes embedding privacy, ELSI, validation, and deployment procedures into ML workflows and governance frameworks.

datagovernance

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the compliance challenges faced by data practitioners in machine learning systems under regulations such as the GDPR and the AI Act, particularly concerning data quality. Through semi-structured interviews with practitioners in the European Union, combined with thematic analysis of regulatory texts and engineering workflows, the research systematically uncovers a structural disconnect between regulation-driven data quality requirements and ML engineering practices. It identifies five core challenges: misalignment between legal principles and engineering implementation, fragmented data pipelines, lack of purpose-built compliance tools, ambiguous accountability, and reactive responses to audits. Building on these findings, the work proposes directions for designing compliance-oriented tooling, establishing effective governance mechanisms, and fostering cultural transformation to bridge the gap between regulatory mandates and practical ML development.

AI Actdata qualityGDPR

This study addresses the compliance challenges faced by machine learning systems in the European Union, where data quality practices often misalign with regulatory requirements due to a lack of actionable guidance. It presents the first systematic mapping of data quality dimensions to specific provisions of EU regulations, resulting in a practical compliance framework. Through an online survey of over 180 European practitioners, combined with regulatory analysis and empirical investigation, the research uncovers a significant gap between current industry practices and regulatory expectations. A key bottleneck identified is insufficient collaboration between technical and legal teams. To bridge this gap, the study advocates for the development of integrated data quality tools and stronger cross-disciplinary collaboration to support the compliant deployment of trustworthy AI systems.

data qualityEU regulationsmachine learning

Lawful and Accountable Personal Data Processing with GDPR-based Access and Usage Control in Distributed Systems

Mar 10, 2025
LT
L. Thomas van Binsbergen
🏛️ University of Amsterdam | Leibniz Institute | TNO

This paper addresses the challenges of GDPR compliance, ambiguous accountability, and high manual auditing costs in cross-organizational distributed data processing. Methodologically, it introduces the first automated compliance framework integrating legal expert judgment with formal reasoning: (i) a purpose-limitation–driven GDPR ontology and semantic model; (ii) a novel deep extension of eFLINT and XACML to support legally precise policy modeling and verifiable enforcement; and (iii) an automated normative reasoning engine that generates auditable legality justifications. Contributions include: transparent and accountable data access control; empirical validation across multiple distributed data space prototypes demonstrating completeness of legal argumentation, accuracy of policy enforcement, and seamless system integrability; and significant reduction in organizational compliance overhead and legal risk.

Automated normative reasoning for GDPR compliance in data processing.Formal ontology and semantics for lawful and accountable data handling.Integration of GDPR-based access and usage control in distributed systems.

Several Issues Regarding Data Governance in AGI

Aug 16, 2025
MH
Masayuki Hatta
🏛️ Surugadai University

This paper addresses unique data governance challenges arising from artificial general intelligence (AGI) systems endowed with recursive self-improvement and self-replication capabilities. It identifies seven urgent, AGI-specific risks surpassing those of conventional AI: autonomous data acquisition bypassing informed consent; data retention decisions driven by optimization objectives rather than human values; supranational, unregulated data sharing among decentralized AGI agents; erosion of data provenance due to dynamic system evolution; ambiguous ownership of AI-generated content; diminished regulatory enforcement across jurisdictions; and rapid obsolescence of static governance frameworks. Employing conceptual analysis and systematic reasoning, the study integrates theories of recursive self-improvement, data provenance, and cross-border regulatory compliance to construct a risk identification and response framework. Its key contribution is a novel tripartite governance paradigm—“embedded safety constraints, real-time adaptive monitoring, and multilateral co-evolution”—advancing data governance from static rule-based models toward continuous, adaptive evolution, thereby offering theoretical foundations and actionable pathways for global AGI policy development.

Addressing autonomous data collection challenges in AGI systemsExamining data retention decisions beyond human control in AGISolving jurisdictional enforcement issues for self-replicating AGI

JustAct+: Justified and Accountable Actions in Policy-Regulated, Multi-Domain Data Processing

Jan 31, 2025
CA
Christopher A. Esterhuyse
🏛️ University of Amsterdam

Ensuring dynamic regulatory compliance in cross-organizational, multi-domain data processing—particularly in healthcare—is challenging due to heterogeneous legal constraints, contractual obligations, and evolving individual consents, requiring simultaneous privacy preservation, accountability distribution, and policy adaptability. Method: We propose a decentralized policy fragmentation framework enabling autonomous agents to generate, propagate, and dynamically compose local policy fragments to justify data operations—without requiring a global policy view. We introduce the first logic-programming–based justifiable action model, integrated with a distributed gossip protocol and externally synchronized configuration. Contribution/Results: The framework achieves reproducible authorization, auditable accountability, and verifiable governance under weak centralization. Evaluated in the Brane healthcare system, it demonstrates robust compliance assurance, seamless cross-domain interoperability, and audit traceability with full reproducibility.

Data GovernancePolicy CompliancePrivacy Protection

Latest Papers

What's happening recently
View more

This study addresses the challenge of reliably formalizing General Data Protection Regulation (GDPR) legal provisions in light of the semantic nuances and context dependence inherent in legal texts. To this end, the authors propose a human-AI collaborative framework that integrates multi-agent large language models with structured human validation. The approach employs a role-differentiated multi-agent system to automatically generate legal scenarios, formalized rules, and atomic facts, while incorporating iterative expert feedback at representational, logical, and legal levels. Empirical results demonstrate that this methodology substantially enhances the accuracy and reliability of legal formalization and yields a high-quality dataset of GDPR formalizations. Crucially, the findings underscore the indispensable role of structured human oversight in managing legal complexity and enabling context-sensitive reasoning.

auto-formalizationGDPRhuman-in-the-loop

This work addresses the compliance challenges in federated data processing arising from heterogeneous cross-organizational access policies, regulatory discrepancies, and long-running workflows. To tackle these issues, the paper proposes a compliance-aware federated data processing framework that uniquely integrates large language models (LLMs) with a “policy-as-code” approach. This integration enables the automatic translation of natural language descriptions of legal and organizational compliance requirements into executable machine-interpretable policies. An orchestration engine then enforces these policies dynamically across end-to-end workflows. Evaluation of the prototype system demonstrates that the proposed method effectively harmonizes multi-source compliance rules, significantly enhancing both compliance assurance and deployment feasibility in federated environments.

Access PoliciesCompliance ManagementFederated Data Processing

This study addresses the challenges posed by divergent and conflicting data protection regulations across jurisdictions, which hinder the early identification of compliance requirements in software development and often lead to costly rework and legal risks. Drawing on interviews with 70 legal experts from G20 and other countries, the research employs systematic content analysis and deductive qualitative methods to distill, for the first time from a legal expert perspective, both commonalities—such as consent—and key divergences—such as the right to be forgotten—across global data protection laws. These insights are innovatively operationalized into a comprehensive set of Data Protection Officer (DPO) user stories mapped to each phase of the software development lifecycle and enterprise architecture layers, significantly enhancing the actionable integration of compliance requirements into early-stage software engineering practices.

data protection regulationsprivacy complianceregulatory data protection requirements

This study addresses the significant challenges in implementing the General Data Protection Regulation (GDPR) rights to rectification and erasure within machine learning (ML) supply chains, particularly in complex, multi-party workflows. It introduces the novel concept of “dark models”—opaque and non-traceable downstream derivative models that exacerbate compliance risks by undermining data subjects’ rights. By integrating legal and artificial intelligence perspectives, the work develops an interdisciplinary analytical framework through a synthesis of scholarly literature and regulatory guidance to evaluate the capacity of current technical approaches to meet GDPR requirements. The analysis reveals that most existing methods fall short of fulfilling core GDPR obligations, identifying critical compliance gaps across the ML supply chain. These findings offer theoretical grounding and actionable pathways for the co-governance of trustworthy AI through aligned institutional and technical solutions.

data subject rightserasureGDPR

This work addresses the limitations of existing data governance tools, which struggle to dynamically adapt to emerging regulations such as India’s Digital Personal Data Protection (DPDP) Act and often lack transparency and explainability, leading to inadequate compliance. To bridge this gap, the paper introduces the first goal-driven agent framework specifically designed for data compliance. The framework integrates a KYU Agent and a Compliance Agent that jointly leverage semantic understanding, user trust modeling, and data sensitivity reasoning, embedding regulatory logic directly into the system to ensure auditable and interpretable decisions. It incorporates anonymization strategies—including masking, pseudonymization, and generalization—and demonstrates significant improvements in DPDP compliance across ten domains, including healthcare, education, and e-commerce, enabling transparent, efficient, and cross-domain adaptive data governance.

compliancedata governanceDPDP Act

Hot Scholars

GL

Guoliang Li

Professor, Tsinghua University
DatabaseBig DataCrowdsourcingData Cleaning & Integration
AX

Amy X. Zhang

Associate Professor, Computer Science & Engineering, University of Washington
social computingHCI
SP

Silvio Peroni

University of Bologna
Semantic PublishingSemantic WebOpen ScienceScience of Science
EW

Eugene Wu

Columbia University
Databasesagent ready systemsdata visualizationdata explanation