Benchmarking AI Performance on End-to-End Data Science Projects

📅 2026-02-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of existing AI benchmarks, which predominantly assess isolated data science capabilities while neglecting systematic evaluation of end-to-end project completion. The authors propose the first comprehensive evaluation framework tailored to full-cycle data science projects, introducing a benchmark comprising 40 real-world tasks that integrate multidimensional competencies—including technical implementation, analytical reasoning, communication, and ethical considerations. They further develop an assessment pipeline combining structured scoring rubrics with automated evaluation procedures. Experimental results demonstrate that state-of-the-art generative AI models perform comparably to junior data scientists on well-structured tasks, yet exhibit substantial performance gaps in tasks requiring subjective judgment, thereby underscoring the continued necessity of human validation in complex data science workflows.

Technology Category

Humans and AI: Other Foundations of Human Computation & AIMachine Learning: Evaluation and AnalysisNatural Language Processing: Generation

Application Category

Economics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAISearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsWeb Mining and Content Analysis: Machine learning and data science for the Web
📝 Abstract
Data science is an integrated workflow of technical, analytical, communication, and ethical skills, but current AI benchmarks focus mostly on constituent parts. We test whether AI models can generate end-to-end data science projects. To do this we create a benchmark of 40 end-to-end data science projects with associated rubric evaluations. We use these to build an automated grading pipeline that systematically evaluates the data science projects produced by generative AI models. We find the extent to which generative AI models can complete end-to-end data science projects varies considerably by model. Most recent models did well on structured tasks, but there were considerable differences on tasks that needed judgment. These findings suggest that while AI models could approximate entry-level data scientists on routine tasks, they require verification.
Problem

Research questions and friction points this paper is trying to address.

AI benchmarking
end-to-end data science
generative AI
automated evaluation
data science workflow
Innovation

Methods, ideas, or system contributions that make the work stand out.

end-to-end benchmarking
generative AI evaluation
automated grading pipeline
data science workflow
AI performance assessment
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
E
Evelyn Hughes
R
Rohan Alexander