PIPES: A Meta-dataset of Machine Learning Pipelines

📅 2025-09-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Addressing key bottlenecks in the Algorithm Selection Problem (ASP)—including high evaluation costs, insufficient pipeline diversity, and severe sampling bias in existing meta-datasets (e.g., OpenML)—this work introduces PIPES, a large-scale, balanced meta-dataset. PIPES encompasses 300 diverse datasets and 9,408 systematically designed machine learning pipelines, spanning broad preprocessing–model combinations. It is the first to achieve comprehensive, bias-mitigated, standardized pipeline metadata collection, fully recording module specifications, execution time, predictions, performance metrics, and failure logs. A unified metadata schema enables deep cross-pipeline and cross-dataset analysis. PIPES substantially enhances reproducibility and scalability in meta-learning research. All code and experimental data are publicly released.

Technology Category

Search and Optimization: Metareasoning and MetaheuristicsMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Ethics — Bias, Fairness, Transparency & Privacy

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsSecurity and Privacy: Data transparency and provenance
📝 Abstract
Solutions to the Algorithm Selection Problem (ASP) in machine learning face the challenge of high computational costs associated with evaluating various algorithms' performances on a given dataset. To mitigate this cost, the meta-learning field can leverage previously executed experiments shared in online repositories such as OpenML. OpenML provides an extensive collection of machine learning experiments. However, an analysis of OpenML's records reveals limitations. It lacks diversity in pipelines, specifically when exploring data preprocessing steps/blocks, such as scaling or imputation, resulting in limited representation. Its experiments are often focused on a few popular techniques within each pipeline block, leading to an imbalanced sample. To overcome the observed limitations of OpenML, we propose PIPES, a collection of experiments involving multiple pipelines designed to represent all combinations of the selected sets of techniques, aiming at diversity and completeness. PIPES stores the results of experiments performed applying 9,408 pipelines to 300 datasets. It includes detailed information on the pipeline blocks, training and testing times, predictions, performances, and the eventual error messages. This comprehensive collection of results allows researchers to perform analyses across diverse and representative pipelines and datasets. PIPES also offers potential for expansion, as additional data and experiments can be incorporated to support the meta-learning community further. The data, code, supplementary material, and all experiments can be found at https://github.com/cynthiamaia/PIPES.git.
Problem

Research questions and friction points this paper is trying to address.

Addressing limited pipeline diversity in OpenML repositories
Overcoming imbalanced algorithm selection in meta-learning datasets
Providing comprehensive pipeline combinations for ASP evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diverse pipeline combinations for completeness
Comprehensive experiment results storage
Expandable meta-dataset for community use
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Cynthia Moreira Maia
Centro de Informática, Universidade Federal de Pernambuco, Recife, Brazil
L
Lucas B. V. de Amorim
Instituto de Computação, Universidade Federal de Alagoas, Maceió, Brazil
G
George D. C. Cavalcanti
Centro de Informática, Universidade Federal de Pernambuco, Recife, Brazil
R
Rafael M. O. Cruz
École de Technologie Supérieure, Université du Québec, Montréal, Canada