🤖 AI Summary
This study addresses the challenge that machine learning task types typically require manual specification and are difficult to infer automatically. To this end, it proposes a novel paradigm for task recognition based on large language models (LLMs), which infers both the data domain and the prediction task solely from target features. The authors construct the first annotated benchmark comprising 625 tabular and time-series datasets and systematically evaluate LLM performance across diverse scenarios. Experimental results demonstrate that the proposed approach achieves an F1 score of 0.98 in tabular settings, outperforming AutoGluon, and attains a best cross-domain F1 score of 0.90. Furthermore, lightweight small models deployed locally achieve an F1 score of 0.75, effectively balancing practical feasibility with predictive accuracy.
📝 Abstract
Machine learning task type identification is essential for constructing valid ML pipelines, yet in practice it is typically specified manually. We investigate whether large language models (LLMs) can infer both the data domain and the downstream prediction task directly from dataset-level information when only the target feature is provided by the user. Together with our LLM-based system we also release an annotated benchmark comprising 625 public tabular and time series datasets. We evaluate the proposed approach in three settings: (i) tabular datasets in comparison with established AutoML heuristics, (ii) cross-domain evaluation across tabular and time series datasets, and (iii) a practical deployment scenario using smaller local models. The results show consistent advantages for LLM-based task type identification, with increasing difficulty in heterogeneous and resource-constrained settings. LLM-based approaches outperform AutoGluon in the tabular setting, reaching 0.98 F1 macro compared to 0.93. In the cross-domain setting, the best model achieves 0.90 F1 macro, while smaller locally deployable models reach 0.75, indicating a trade-off between deployment feasibility and accuracy.