Using Small Language Models to Reverse-Engineer Machine Learning Pipelines Structures

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of extracting machine learning pipeline stages, which is constrained by domain diversity and where existing methods rely on manual annotation or limited classifiers. This work systematically investigates, for the first time, the potential of small language models (SLMs) to parse ML pipeline structures leveraging their inherent code comprehension capabilities without fine-tuning, employing Cochran’s Q test, McNemar’s test, and goodness-of-fit evaluations for rigorous assessment. The findings indicate that while SLMs demonstrate robust performance, they do not surpass existing classifiers; however, the core contribution lies in revealing that different classification approaches significantly influence practical insights. Despite the limitation of high inference costs, this research establishes a novel paradigm for automated ML structure parsing.
📝 Abstract
Context: Once defined a taxonomy of stages structuring Machine Learning (ML) pipelines (e.g. Data Preprocessing, Modeling...), extracting these stages from source code is key for better understanding ML practices. However, the diversity caused by the constant evolution of ML (e.g., algorithms, datasets) makes this task challenging. Existing approaches either rely on non-scalable manual labeling or on classifiers that do not properly support domain's diversity. These limitations call for more reliable solutions. Objective: We evaluate whether Small Language Models (SLMs) can leverage their code understanding and classification abilities to address these limitations, and enhance our understanding of practices in ML. Method: We conduct a confirmatory study based on two relevant reference works representing current limitations in the state-of-the-art. We first compare several SLMs using Cochran's Q test, then evaluate the best-performing model against reference studies via two McNemar's tests. An additional Cochran's Q test examines how taxonomy definition variations affect the SLM performance. Finally, goodness-of-fit tests compare ML practice insights from SLM classification with those from prior studies. Results: First, we found that the taxonomy wording significantly impacts classification performance. Second, the best performing SLM yielded good results, yet, without outperforming other classifiers. Third, the three classification methods led to significantly different insights, with varying effect sizes, when exploring practices of data scientists. Conclusions: Limitations of existing classification methods bias our understanding of ML practices. While current SLMs show promising results without prior fine-tuning, they still exhibit common limitations, in addition to inference high costs challenging their applicability in large-scale studies.
Problem

Research questions and friction points this paper is trying to address.

Machine Learning Pipelines
Reverse Engineering
Small Language Models
Code Classification
Taxonomy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Small Language Models
Machine Learning Pipelines
Reverse Engineering
Code Classification
Taxonomy