ModelLakeFishing: Efficient Retrieval over Million-Scale Model Lakes

πŸ“… 2026-10-04
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the prohibitive cost of retrieving models adapted to new datasets from million-scale open model lakes. We propose a millisecond-level retrieval framework that integrates sparse relational evidence graph encoding with hierarchical navigable search. Specifically, our method constructs relational evidence graphs and employs a relation-aware graph neural network to learn model embeddings. These embeddings are combined with Hierarchical Navigable Small World (HNSW) indexing for rapid recall, followed by task-prior-guided reranking. As the first efficient retrieval solution tailored for large-scale model lakes, our approach achieves a Gold@10 of 0.2968, retaining 93.47% of the performance of an exhaustive baseline while requiring only 0.747 milliseconds in median retrieval time, thereby demonstrating an exceptional balance between accuracy and efficiency.
πŸ“ Abstract
Open model lakes may contain millions of reusable models, making it costly to identify suitable models for a new dataset. We present ModelLakeFishing, a model-retrieval framework for queries specifying a target dataset, prediction task, and evaluation metric. It consolidates metadata and historical evaluations into a model-dataset-task evidence graph, learns model and query embeddings with a relation-aware graph encoder, and indexes model embeddings using Hierarchical Navigable Small World (HNSW) search. At query time, HNSW retrieves 1,000 candidates without scoring every model, after which a training-side task prior reranks candidates for the requested metric and returns the top 10. We evaluate on a lake of 3,016,439 models and 247,803 observed model-dataset performance pairs using three root-aware splits that hold test performance edges out of representation learning and retrieval. ModelLakeFishing achieves a mean eligible-query gold@10 of 0.2968, recovering the observed-best model in the top 10 for 29.68% of eligible queries and retaining 93.47% of the gold@10 of an exhaustive baseline using the same scoring and reranking procedure. Given precomputed query embeddings, retrieval and reranking take 0.747 ms median and 1.102 ms at the 95th percentile. These results demonstrate efficient retrieval over million-model lakes from sparse relational evidence.
Problem

Research questions and friction points this paper is trying to address.

Model Retrieval
Model Lake
Efficient Search
Reusable Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Model Retrieval
Evidence Graph
Relation-Aware Graph Encoder
HNSW Search
Task Priorior Reranking
πŸ”Ž Similar Papers
2024-03-04arXiv.orgCitations: 0