design data governance

Design data governance is the practice of designing, building, and analysing the organisational policies, roles, processes, and technical controls that govern how data and datasets are collected, accessed, used, shared, and retired. It produces governance frameworks and artefacts such as dataset-specific access and use policies, role-based and contractual safeguards, consent and traceability mechanisms, audit and compliance documentation, governance-driven platform features, human oversight arrangements, socio-technical and participatory decision-making processes, and validity-based governance instruments to enable accountability and community-centred oversight.

designdatagovernance

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$223K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the limitations of current data governance frameworks, which overly emphasize compliance and risk mitigation at the expense of enabling responsible cross-organizational data reuse for public benefit. To overcome this inward-looking paradigm, the paper proposes a novel institutional function—strategic data stewardship—centered on ecosystem collaboration and public value creation. It introduces an actionable Data Stewardship Canvas to operationalize this approach, integrating institutional design, governance principles, and practical mechanisms. The framework articulates core principles, roles, and capability models tailored to support real-world implementations in data collaboratives, data spaces, and data commons. By doing so, it aims to establish trustworthy, lawful, and efficient pathways for data reuse in the AI era, effectively bridging the gap between data availability and actual accessibility.

AI ethicsdata accessibilitydata governance

Toward Effective AI Governance: A Review of Principles

May 29, 2025
DR
Danilo Ribeiro
🏛️ Zup Innovation

Current AI governance research lacks systematic integration of diverse frameworks and practices, with notable gaps in the operationalizability of key mechanisms and the implementation of inclusive, stakeholder-centered approaches. To address this, we conduct a rapid three-tier literature review, systematically synthesizing nine authoritative IEEE/ACM reviews published between 2020 and 2024. We introduce the novel “thematic semantic synthesis” analytical paradigm to identify high-frequency governance frameworks (e.g., the EU AI Act, NIST AI Risk Management Framework), core principles (e.g., transparency, accountability), and stakeholder role distributions. Our analysis reveals four critical knowledge gaps in AI governance scholarship and practice. Based on these findings, we propose a rigorously grounded, organizationally feasible governance roadmap—bridging theoretical advancement and real-world implementation. This work contributes both empirical evidence and methodological innovation to advance AI governance research and practice.

Addressing gaps in empirical validation and inclusivityIdentifying key principles like transparency and accountabilitySynthesizing diverse AI governance frameworks and practices

Several Issues Regarding Data Governance in AGI

Aug 16, 2025
MH
Masayuki Hatta
🏛️ Surugadai University

This paper addresses unique data governance challenges arising from artificial general intelligence (AGI) systems endowed with recursive self-improvement and self-replication capabilities. It identifies seven urgent, AGI-specific risks surpassing those of conventional AI: autonomous data acquisition bypassing informed consent; data retention decisions driven by optimization objectives rather than human values; supranational, unregulated data sharing among decentralized AGI agents; erosion of data provenance due to dynamic system evolution; ambiguous ownership of AI-generated content; diminished regulatory enforcement across jurisdictions; and rapid obsolescence of static governance frameworks. Employing conceptual analysis and systematic reasoning, the study integrates theories of recursive self-improvement, data provenance, and cross-border regulatory compliance to construct a risk identification and response framework. Its key contribution is a novel tripartite governance paradigm—“embedded safety constraints, real-time adaptive monitoring, and multilateral co-evolution”—advancing data governance from static rule-based models toward continuous, adaptive evolution, thereby offering theoretical foundations and actionable pathways for global AGI policy development.

Addressing autonomous data collection challenges in AGI systemsExamining data retention decisions beyond human control in AGISolving jurisdictional enforcement issues for self-replicating AGI

Towards Data Governance of Frontier AI Models

Dec 05, 2024
JH
Jason Hausenloy
🏛️ University of California, Berkeley | University of Cambridge | Independent

The rapid scaling of frontier AI models has introduced novel societal and technical risks, necessitating governance frameworks that move beyond model-centric regulation. Method: This paper proposes “frontier data governance” as a paradigm shift—treating data not merely as a risk vector but as a primary governance lever. We systematically analyze stakeholders across the AI data supply chain and synthesize 15 existing governance mechanisms. Building on this, we introduce five underexplored yet actionable policy instruments: data watermarking (canary tokens), automated content filtering, mandatory dataset registration, generative algorithm security hardening, and Know-Your-Data-Provider (KYD) requirements. Through empirical integration of data provenance, secure data encapsulation, and synthetic data oversight, we design and validate five implementable policy proposals. Contribution/Results: This work delivers the first comprehensive, data supply chain–oriented governance roadmap for frontier AI, offering both theoretical foundations and operational blueprints to inform global AI regulatory practice.

Addressing challenges from data's non-rival and replicable natureHow data enables governance for frontier AI modelsProposing policy mechanisms for data supply chain actors

Artificial Intelligence Governance For Businesses

Nov 20, 2020
JS
Johannes Schneider
🏛️ University of Liechtenstein | Ruhr University of Bochum

Existing AI governance research predominantly addresses macro-level regulatory principles, leaving a critical gap in enterprise-level implementation frameworks. This paper proposes a three-layer conceptual framework for organizational AI governance—spanning data, models, and systems—structured around the triad of “actor–artifact–mechanism.” It introduces two novel elements: (1) a data-value quantification methodology and (2) formally defined, role-specific AI governance positions. Leveraging literature-driven modeling, multi-dimensional structural decomposition, and cross-layer alignment techniques, the framework is designed for seamless integration into existing corporate governance infrastructures. The resulting implementation pathway bridges the translational gap between high-level regulatory guidance and operational AI governance practice. By unifying theoretical rigor with practical feasibility, this work establishes a new paradigm for institutionalized AI governance, directly addressing the longstanding challenge of converting abstract governance principles into actionable, organizationally embedded practices. (149 words)

Addressing lack of AI governance frameworks for businessesBridging gaps between academic research and practical implementationIntegrating data, ML models, and AI systems governance

Latest Papers

What's happening recently
View more

This study addresses the structural challenges and systemic risks confronting data governance in the AI era—particularly concerning access, reuse, and sovereignty. Integrating perspectives from AI governance, digital public infrastructure, and geopolitics, the research pioneers a systematic application of horizon-scanning–based qualitative signal detection. Through two rounds of expert forecasting workshops and thematic clustering analysis, it identifies seven key trends and their reinforcing feedback mechanisms. The findings underscore the inevitable convergence of data and AI governance and articulate a forward-looking intervention agenda aimed at shaping institutional trajectories before path dependencies become entrenched. The work offers policymakers a structured diagnostic framework and an adaptive governance blueprint to navigate emerging complexities in data-driven societies.

AI governanceanticipatory governancedata governance

This study addresses the prevailing overemphasis on legal compliance in student data governance within learning analytics, which often neglects critical ethical dimensions such as fairness, student autonomy, accountability, and educational purpose. To bridge this gap, the work proposes LEAGUE—a six-pillar ethical governance framework encompassing Legitimacy, Equity, Autonomy, Governance, Utility, and Ethics-by-Design—integrating insights from learning analytics, data ethics, and capability-oriented theories of educational justice. Moving beyond conventional compliance paradigms, this framework pioneers the application of value-sensitive design in the field. Through conceptual review, interdisciplinary theoretical synthesis, and a case analysis of an early warning system, the study demonstrates the framework’s feasibility in enhancing transparency, educational meaningfulness, and ethical justifiability, offering a theoretically grounded yet practically actionable pathway for ethically robust learning analytics.

data ethicseducational justiceethical governance

This work addresses the fragmentation of governance evidence in automated decision systems caused by heterogeneous log formats by proposing the Decision Event Schema (DES)—a unified tracing specification based on JSON Schema. DES uniquely integrates four infrastructure layers: machine learning inference, rule evaluation, cross-system coupling, and governance metadata. It introduces degradation-resistant field design and a three-tier evidence strategy—lightweight, sampled, and complete—to accommodate varying risk profiles and throughput requirements. Among over 25 existing logging formats, DES is the only one that comprehensively covers all four layers. Empirical validation demonstrates its compatibility with high-throughput production environments, offering practitioners a readily adoptable or extensible reference standard and enabling regulators to map compliance requirements to minimal evidence levels.

decision loggingFragmented Trace Problemgovernance evidence

In multi-stakeholder platforms, software architecture decisions often implicitly entrench conflicting requirements without systematic support for mapping governance principles to technical design. This work proposes the first governance-architecture alignment framework, explicitly linking five core governance principles to the space of architectural decisions, thereby rendering implicit governance stances identifiable and contestable. The framework also exposes how default technical choices can obscure underlying value commitments. Feasibility is preliminarily demonstrated through a constructive case study of a pig-farming knowledge platform in Rwanda. Future work will employ pre- and post-intervention user judgment studies to evaluate the framework’s impact on actual governance outcomes.

architectural decisionsgovernancemulti-stakeholder platforms

This study addresses the strategic lock-in and operational risks organizations face when relying on commercial data intermediaries to ensure the timeliness and reliability of master data. It pioneers the systematic integration of Self-Sovereign Identity (SSI) into master data management by synthesizing insights from hermeneutic literature review, expert interviews, and design science research methodologies. The resulting design theory embeds a trustworthy master data management framework within a data space reference architecture, emphasizing data sovereignty, reliability, and accountability. Validated through evaluation by industry experts, the proposed framework enables trusted, controllable data sharing and governance within data ecosystems.

data qualitydata sovereigntymaster data management

Hot Scholars

RB

Rishi Bommasani

CS PhD, Stanford University
Societal Impact of AIAI PolicyAI GovernanceFoundation Models
DS

Dawn Song

Professor of Computer Science, UC Berkeley
Computer Security and Privacy
AK

Atoosa Kasirzadeh

Carnegie Mellon University
AI EthicsAI GovernancePhilosophyMathematical Optimization
RM

Rashid Mushkani

University of Montreal I Mila
Public (Space & Life)Sociotechnical AIUrban AnalyticsCommunity-Centered AI
SC

Stephen Casper

PhD student, MIT
AI safetyAI responsibilityred-teamingrobustness