privacy and ethics compliance

Designing data collection, labeling, and curation workflows that adhere to privacy, consent, age verification, and broader ethical safeguards for sensitive datasets. This includes procedures for assembling corpora (faces, neonatal videos, social media, touch clips) while ensuring correct labels, minimising harm, and documenting ethical decisions.

privacyandethicscompliance

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Neglected Risks: The Disturbing Reality of Children's Images in Datasets and the Urgent Call for Accountability

Apr 20, 2025
CC
Carlos Caetano
🏛️ Universidade Estadual de Campinas (UNICAMP) | Universidade Federal de Minas Gerais (UFMG) | Universidade de São Paulo (USP) | Instituto Alana | University of Sheffield

This work exposes systemic ethical risks—including privacy violations, lack of informed consent, and absence of accountability—arising from the pervasive presence of children’s images in AI training datasets. Addressing the challenges of detecting and filtering such images in existing large-scale datasets (e.g., Open Images V7), we propose the first open-source, reproducible, ethics-first detection-and-removal pipeline. Our approach integrates fine-tuned vision-language models, facial analysis, age estimation, and context-aware semantic filtering to robustly identify child-associated content. Evaluated on the #PraCegoVer benchmark and an Open Images subset, it achieves high recall—effectively identifying child-related images among 100,000 samples. Empirical validation demonstrates its practical utility in data purification for visual question answering (VQA) tasks. Beyond technical contribution, this work advances normative frameworks for child data governance and catalyzes a cross-institutional ethical action initiative.

Ethical concerns of children's images in AI datasetsLack of methods to detect and remove children's imagesNeed for tools to protect children's privacy rights

Public vision datasets (e.g., ImageNet, COCO, CelebA) pose significant privacy leakage and algorithmic bias risks in high-stakes domains such as healthcare and security. Method: This work introduces the first end-to-end computer vision ethics assessment framework—spanning data acquisition, model training, and deployment—and proposes a dual-track validation standard integrating privacy compliance and bias quantification. The framework unifies data provenance analysis, statistical bias detection, anonymization efficacy evaluation, and transparency auditing, supported by cross-dataset ethical benchmarking experiments. Contribution/Results: Empirical evaluation reveals substantial identity re-identification risks and demographic representation imbalances across seven widely used datasets. Based on these findings, we propose three actionable data governance protocols—formally adopted by the IEEE P7003 Ethics in Action Standard Working Group—to mitigate ethical harms while ensuring technical feasibility and regulatory alignment.

Addressing privacy concerns in computer vision datasetsDeveloping ethical frameworks for responsible AI developmentMitigating bias in publicly available visual datasets

How do data owners say no? A case study of data consent mechanisms in web-scraped vision-language AI training datasets

Nov 10, 2025
CP
Chung Peng Lee
🏛️ Princeton University | University of Washington | Carnegie Mellon University | University of Toronto

This study addresses the critical challenge of respecting data owners’ intent and ensuring copyright compliance in training AI vision-language models. Methodologically, it integrates sample-level signals (e.g., copyright notices, watermarks, metadata) with domain-level policies (e.g., robots.txt, terms of service), combining statistical estimation and content detection to systematically characterize multi-channel opt-out signals—marking the first such comprehensive analysis. Key findings reveal: (i) at least 122 million samples in CommonPool contain explicit copyright indicators; (ii) 60% of top-domain training data originates from websites explicitly prohibiting web crawling; and (iii) 9–13% of images contain watermarks, yet current detection methods exhibit high false-negative rates. These results expose substantial gaps in AI data pipelines’ ability to recognize and respond to refusal signals. Accordingly, the study proposes a unified data consent framework tailored for AI training, offering an empirically grounded foundation and actionable pathway toward accountable, auditable, and legally compliant data governance.

Analyzing copyright notices, watermarks and ToS compliance in web datasetsInvestigating how data owners express consent for AI data scraping and trainingRevealing current AI data collection fails to respect owner consent mechanisms

Ethical documentation for multimodal datasets is widely absent or inconsistent, severely impeding the development of responsible AI. To address this, we propose TEDI, the first standardized ethical assessment framework comprising 143 fine-grained, verifiable metrics across core dimensions—including informed consent, privacy protection, and harmful content mitigation. Leveraging human annotation and metadata modeling, we empirically evaluate over 100 speech-related multimodal datasets. Our analysis reveals that web-crawled datasets exhibit significantly poorer ethical documentation quality compared to crowdsourced or directly collected ones, uncovering a strong correlation between data acquisition methodology and documentation rigor. Beyond diagnostic insights, TEDI establishes the first reproducible, extensible benchmark for ethical completeness—enabling automated documentation parsing and supporting trustworthy AI governance practices.

Difficulty in comparing ethical indicators across multimodal datasetsLack of insights into trustworthy and ethical dataset attributesLimited documentation on consent, privacy, and harmful content in datasets

This study examines the ethical and legal challenges arising from AI-driven data collection during 2023–2024, identifying core risks—including absent informed consent, amplified algorithmic bias, and systemic privacy erosion—across healthcare, finance, and smart city domains. It comparatively analyzes regulatory approaches in the EU, U.S., and China. Methodologically, it integrates policy text analysis, cross-jurisdictional compliance mapping, empirical case studies, and multi-stakeholder Delphi consultation. The study proposes a three-dimensional adaptive governance framework comprising *legal alignment*, *technical safeguards* (e.g., embedded differential privacy), and *dynamic ethical assessment*. Its key contributions include advancing context-sensitive regulation and fostering transnational standardization for AI data governance; it delivers an actionable AI data governance roadmap, already adopted by three international digital ethics working groups and informing the design of two regional AI regulatory pilot programs.

Compares regulatory approaches in EU, US, and China for AI governance.Examines ethical and legal implications of AI-driven data collection.Proposes solutions for balancing AI innovation with privacy protection.

Latest Papers

What's happening recently
View more

This study addresses the critical privacy risks posed by sensitive medical images—such as prenatal ultrasound scans containing personally identifiable information like names and locations—within large-scale public image datasets like LAION-400M. It presents the first systematic investigation into the presence of such high-risk content in general-purpose datasets used for training generative models. By leveraging CLIP embedding similarity search, image content analysis, and named entity recognition, the authors successfully retrieved thousands of ultrasound images from LAION-400M that expose identifiable patient information. The findings demonstrate the widespread inclusion of sensitive medical data in training corpora, highlighting significant privacy vulnerabilities. Based on this empirical evidence, the work proposes concrete recommendations for privacy-preserving dataset curation and usage, offering a foundation for responsible data governance in AI development.

image datasetspregnancy ultrasoundprivacy

This study addresses the “protection paradox” wherein AI-driven data analytics, while intended to safeguard vulnerable populations, may inadvertently exacerbate their vulnerability through inherent technical processes. Conceptualizing vulnerability as a dynamically constructed outcome of data practices, the work innovatively integrates ethical considerations into four critical stages of the AI pipeline: dataset design, operationalization and modeling, inferential logic, and dissemination strategies. Employing AI for Social Good (AI4SG) methodologies—including computer vision, critical dataset analysis, and inference auditing—the research identifies four key factors that contribute to algorithmic fragility. Building on these insights, the authors propose a reflexive ethical roadmap that enables researchers to navigate platform-based data studies while mitigating risks of computational exposure and exploitation stemming from well-intentioned interventions.

AI ethicsdata practicesplatformized data

This study addresses the lack of verifiable ethical safeguards in existing AI-driven digital phenotyping systems that leverage financial behavior data for mental health assessment, which are vulnerable to risks concerning informed consent, privacy, and fairness. To bridge this gap, the authors propose a computational ethics framework that formalizes ethical principles as deontic temporal logic constraints. By integrating the Z3 SMT solver, the framework enables machine-verifiable, real-time compliance checks. It further incorporates an ethics agent and counterexample generation mechanism to dynamically monitor and detect violations. Evaluated in a finance–mental health use case, the framework successfully eliminates scenarios violating specified ethical properties, thereby transcending conventional static, post-hoc compliance documentation and establishing a foundation for continuous, auditable ethical assurance in AI systems.

AI ethicsdigital phenotypingethical governance

Assessing metadata privacy in neuroimaging

Sep 18, 2025
EK
Emilie Kibsgaard
🏛️ Copenhagen University Hospital | Instituto de Física de Cantabria | Stanford University

This study addresses metadata privacy risks in neuroimaging data sharing by systematically evaluating re-identification vulnerabilities in publicly available BIDS-formatted datasets. We developed and applied metaprivBIDS—a novel tool enabling the first automated, standardized privacy audit of tabular metadata (e.g., demographics, clinical scores)—integrating statistical and semantic analyses to detect cross-population differences in de-identification efficacy. Results indicate low re-identification risk for clinical scores, whereas demographic variables—including age, sex, and nationality—constitute the primary privacy bottleneck. While most datasets exhibit no critical vulnerabilities, widespread mild information leakage persists and remains exploitable. Based on these findings, we propose a tiered mitigation strategy. This work establishes a reproducible, scalable privacy assessment framework for neuroscientific data governance, grounded in empirical evidence and aligned with FAIR and GDPR principles.

Assessing metadata privacy risks in neuroimaging data sharingDeveloping practical measures for safer neuroimaging data sharingEvaluating reidentification vulnerabilities from demographic and clinical data

Hot Scholars

JS

Jose Such

Research Professor, Spanish National Research Council (CSIC)
Privacy & SecurityArtificial IntelligenceHuman-Computer Interaction
SZ

Shuning Zhang

Tsinghua University
HCIUsable Privacy and SecurityAI
TL

Tianshi Li

Assistant Professor, Northeastern University
Human-Computer InteractionPrivacyHuman-Centered AI Privacy
CS

Cristiana Santos

Utrecht University
Compliance with Data Protection LawDark PatternsTracking
RA

Ruba Abu-Salma

Senior Lecturer (~Associate Professor) in Computer Science, King’s College London
Cybersecurityprivacy-enhancing technologies (PETs)HCIusable security and privacy