🤖 AI Summary
This study addresses the limitation of existing chart digitization benchmarks, which rely on synthetic data and lack paired validation between real scientific figures and their source data. To overcome this, we propose an automated pipeline that constructs a benchmark from real biological preprints and their author-released source datasets. By employing mapping algorithms to identify reconstructable panels alongside human verification for generating quantitative questions, our approach achieves precise figure-data alignment and evaluation. We introduce PlotGround-1k, a novel benchmark demonstrating that state-of-the-art multimodal models attain 87.5% accuracy under a ±5% tolerance. Furthermore, providing source tables enables coding agents to reach 97.4% accuracy while reducing computational costs by 72%.
📝 Abstract
Scientific figures often encode quantitative results that are not readily available in machine-readable form, making accurate plot digitization important for verifying and reusing published findings. Yet it remains unclear how accurately current models recover plotted values from real scientific figures, as existing benchmarks rely largely on synthetic charts or cover only a limited range of chart types. We introduce PlotGround, an automated pipeline for building plot digitization benchmarks from real scientific figures and their author-released source data. PlotGround maps figures to source tables, identifies reconstructable panels, and generates quantitative questions with source-grounded reference values. We use PlotGround to construct PlotGround-1k, a human-verified benchmark of 1,119 questions from 1,066 bioRxiv preprints. Across sixteen multimodal models, the best reaches 87.5% accuracy at a $\pm 5\%$ relative-error tolerance. Tightening the tolerance to $\pm 2\%$ lowers every model's accuracy by 11-24 percentage points, revealing a gap between approximate visual reading and precise quantitative recovery. PlotGround's paired figure-source structure lets us compare how accurately the same values are recovered from figures and from source tables. Providing source tables instead of figures raises a coding agent's accuracy from 90.0% to 97.4% while cutting cost by 72%.