🤖 AI Summary
This work addresses the challenge software engineers face in efficiently identifying suitable pre-trained models and datasets for software engineering tasks amid the vast landscape of machine learning assets. To bridge this gap, we propose and implement MLAssetSelection—the first asset selection tool tailored specifically for the software engineering domain. By automatically harvesting relevant assets from platforms such as Hugging Face and integrating multidimensional evaluation metrics, a configurable leaderboard, requirement-based filtering mechanisms, and personalized recommendations, the system enables real-time updates and user customization. Empirical evaluation demonstrates that MLAssetSelection substantially improves the efficiency of model and dataset selection, thereby filling a critical void in domain-specific intelligent asset curation tools.
📝 Abstract
The rapid growth of machine learning assets has made it increasingly difficult for software engineers to identify models and datasets that match their specific needs. Browsing large registries, such as Hugging Face, is time-consuming, error-prone, and rarely tailored to Software Engineering (SE) tasks. We present MLAssetSelection, a web application that automatically extracts SE assets and supports four key functionalities: (i) a configurable leaderboard for ranking models across multiple benchmarks and metrics; (ii) requirements-based selection of models and datasets; (iii) real-time automated updates through scheduled jobs that keep asset information current; and (iv) user-centric features including login, personalized asset lists, and configurable alert notifications. A demonstration video is available at https://youtu.be/t6CJ6P9asV4.