🤖 AI Summary
This study addresses the time-consuming and error-prone nature of manual metadata extraction for machine learning datasets by constructing the first end-to-end evaluation benchmark based on the Croissant community standard, encompassing 602 papers. Furthermore, it proposes a two-tier automated evaluation framework that integrates human auditing with large language model (LLM) judges. Through systematic comparisons of state-of-the-art LLMs, open-source models, and agent-based architectures, our findings reveal that single-pass full-context extraction significantly outperforms decomposed agent frameworks, particularly in processing long-text responsible AI fields. All benchmark data and code resources associated with this project have been made publicly available.
📝 Abstract
Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. The benchmark comprises 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields. Using this benchmark, we evaluate a range of extraction systems spanning frontier models, open-weight models, and agentic architectures, under a two-tier evaluation framework that combines rule-based scoring with an LLM judge selected via human audit. We find that single-pass extraction consistently outperforms the four agentic architectures we evaluate: across backbones, these decomposed variants achieve lower accuracy than a single full-context pass. The largest gap appears on long-form RAI fields, which require synthesizing and interpreting information scattered across a paper rather than copying it from a single location, a setting where current systems remain far from reliable. We release the benchmark, evaluation code, judge audit, a live demo, and a leaderboard open to new systems.