A Tool for Automatically Cataloguing and Selecting Pre-Trained Models and Datasets for Software Engineering

📅 2026-01-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge software engineers face in efficiently identifying suitable pre-trained models and datasets for software engineering tasks amid the vast landscape of machine learning assets. To bridge this gap, we propose and implement MLAssetSelection—the first asset selection tool tailored specifically for the software engineering domain. By automatically harvesting relevant assets from platforms such as Hugging Face and integrating multidimensional evaluation metrics, a configurable leaderboard, requirement-based filtering mechanisms, and personalized recommendations, the system enables real-time updates and user customization. Empirical evaluation demonstrates that MLAssetSelection substantially improves the efficiency of model and dataset selection, thereby filling a critical void in domain-specific intelligent asset curation tools.

Technology Category

Machine Learning: Hardware-aware MLSearch and Optimization: Metareasoning and MetaheuristicsApplication Domains: Software Engineering

Application Category

User Modeling, Personalization and Recommendation: ML for personalized search and recommendationsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
The rapid growth of machine learning assets has made it increasingly difficult for software engineers to identify models and datasets that match their specific needs. Browsing large registries, such as Hugging Face, is time-consuming, error-prone, and rarely tailored to Software Engineering (SE) tasks. We present MLAssetSelection, a web application that automatically extracts SE assets and supports four key functionalities: (i) a configurable leaderboard for ranking models across multiple benchmarks and metrics; (ii) requirements-based selection of models and datasets; (iii) real-time automated updates through scheduled jobs that keep asset information current; and (iv) user-centric features including login, personalized asset lists, and configurable alert notifications. A demonstration video is available at https://youtu.be/t6CJ6P9asV4.
Problem

Research questions and friction points this paper is trying to address.

pre-trained models
datasets
software engineering
model selection
asset cataloguing
Innovation

Methods, ideas, or system contributions that make the work stand out.

pre-trained model selection
software engineering datasets
automated asset cataloguing
requirements-based recommendation
model leaderboard
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Alexandra González
Universitat Politècnica de Catalunya, Barcelona, Spain
O
Oscar Cerezo
Universitat Politècnica de Catalunya, Barcelona, Spain
X
Xavier Franch
Universitat Politècnica de Catalunya, Barcelona, Spain
S
Silverio Martínez-Fernández
Universitat Politècnica de Catalunya, Barcelona, Spain