CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the time-consuming and error-prone nature of manual metadata extraction for machine learning datasets by constructing the first end-to-end evaluation benchmark based on the Croissant community standard, encompassing 602 papers. Furthermore, it proposes a two-tier automated evaluation framework that integrates human auditing with large language model (LLM) judges. Through systematic comparisons of state-of-the-art LLMs, open-source models, and agent-based architectures, our findings reveal that single-pass full-context extraction significantly outperforms decomposed agent frameworks, particularly in processing long-text responsible AI fields. All benchmark data and code resources associated with this project have been made publicly available.
📝 Abstract
Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. The benchmark comprises 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields. Using this benchmark, we evaluate a range of extraction systems spanning frontier models, open-weight models, and agentic architectures, under a two-tier evaluation framework that combines rule-based scoring with an LLM judge selected via human audit. We find that single-pass extraction consistently outperforms the four agentic architectures we evaluate: across backbones, these decomposed variants achieve lower accuracy than a single full-context pass. The largest gap appears on long-form RAI fields, which require synthesizing and interpreting information scattered across a paper rather than copying it from a single location, a setting where current systems remain far from reliable. We release the benchmark, evaluation code, judge audit, a live demo, and a leaderboard open to new systems.
Problem

Research questions and friction points this paper is trying to address.

Croissant metadata
metadata extraction
ML datasets
benchmark evaluation
Responsible AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Croissant Metadata
Metadata Extraction Benchmark
Agentic Architectures
Two-tier Evaluation Framework
Responsible AI
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Berke Arda
ETH Zurich
A
Ahmetcan Yavuz
ETH Zurich
P
Paul Gerry
CSAIL, MIT
S
Sebastian Lobentanzer
Helmholtz Zentrum München
N
Nobin Sarwar
University of Maryland, Baltimore County
J
Joan Giner-Miguelez
Barcelona Supercomputing Center
K
Kongtao Chen
Google
Luyao Zhang
Luyao Zhang
Duke Kunshan University
algorithmic game theorymechanism designmachine learningblockchainexplainable AI
Mrinmaya Sachan
Mrinmaya Sachan
Assistant Professor, ETH Zürich
Natural Language ProcessingReasoningAI for Education
Mubashara Akhtar
Mubashara Akhtar
ETH AI Center fellow at ETH Zurich
NLPMultimodalityBenchmarking & Evaluation