A Critical Field Guide for Working with Machine Learning Datasets

📅 2025-01-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Machine learning datasets pose multifaceted risks concerning technical reliability, legal compliance, and ethical legitimacy. Method: This paper introduces the first systematic, full-lifecycle dataset governance framework, integrating critical AI theory with applied data science. It employs interdisciplinary practices—including data auditing, provenance analysis, bias detection, regulatory compliance assessment, and participatory workshops—to operationalize abstract ethical principles without reliance on specific algorithms or tools. Contribution/Results: The work establishes the first generalizable governance paradigm that enables concurrent technical, legal, and ethical reflection—bridging the gap between pragmatic guidance and humanistic critique. Its open-source guidelines have been widely adopted by educators, media organizations, and open-source communities, significantly enhancing practitioners’ awareness of latent dataset risks. Moreover, the framework has directly catalyzed the publication of accountability statements and usage constraints by multiple major public datasets.

Technology Category

Machine Learning: Ethics, Bias, and FairnessPhilosophy and Ethics of AI: AI & Law, Justice, Regulation & GovernanceNatural Language Processing: Ethics — Bias, Fairness, Transparency & Privacy

Application Category

Security and Privacy: Data transparency and provenanceResponsible Web: Ethical and legal aspects of web-scale data analysis, uses and collection practicesSemantics and Knowledge: Provenance, trust, security and privacy, and ethical issues in managing semantic data
📝 Abstract
Machine learning datasets are powerful but unwieldy. Despite the fact that large datasets commonly contain problematic material--whether from a technical, legal, or ethical perspective--datasets are valuable resources when handled carefully and critically. A Critical Field Guide for Working with Machine Learning Datasets suggests practical guidance for conscientious dataset stewardship. It offers questions, suggestions, strategies, and resources for working with existing machine learning datasets at every phase of their lifecycle. It combines critical AI theories and applied data science concepts, explained in accessible language. Equipped with this understanding, students, journalists, artists, researchers, and developers can be more capable of avoiding the problems unique to datasets. They can also construct more reliable, robust solutions, or even explore new ways of thinking with machine learning datasets that are more critical and conscientious.
Problem

Research questions and friction points this paper is trying to address.

Responsible AI
Machine Learning Ethics
Data Utilization Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Responsible AI
Ethical Data Usage
Machine Learning Reliability
🔎 Similar Papers
No similar papers found.
S
Sarah Ciston
M
Mike Ananny
K
Kate Crawford