Benchmarking Few-shot Transferability of Pre-trained Models with Improved Evaluation Protocols

📅 2026-02-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of a unified, rigorous, and realistic evaluation protocol in few-shot transfer learning, which has led to unreliable method comparisons. To this end, we introduce the FEWTRANS benchmark—comprising ten diverse datasets—and the Hyperparameter Ensemble (HPE) evaluation protocol, which effectively mitigates validation set hallucination under data scarcity. Using this framework, we systematically demonstrate for the first time that the choice of pretrained model is more critical than the complexity of the transfer algorithm. We also quantify the performance collapse of multimodal models in specialized domains due to linguistic rarity. Our analysis reveals that simple full-parameter fine-tuning consistently outperforms most sophisticated methods, owing to its ability to flexibly reshape distributed representations and high-level semantic features. The FEWTRANS benchmark is publicly released to provide the community with a reproducible evaluation standard.

Technology Category

Machine Learning: Transfer, Domain Adaptation, Multi-Task LearningNatural Language Processing: Language Grounding & Multi-modal NLPComputer Vision: Multi-modal Vision

Application Category

Web Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Few-shot transfer has been revolutionized by stronger pre-trained models and improved adaptation algorithms.However, there lacks a unified, rigorous evaluation protocol that is both challenging and realistic for real-world usage. In this work, we establish FEWTRANS, a comprehensive benchmark containing 10 diverse datasets, and propose the Hyperparameter Ensemble (HPE) protocol to overcome the "validation set illusion" in data-scarce regimes. Our empirical findings demonstrate that the choice of pre-trained model is the dominant factor for performance, while many sophisticated transfer methods offer negligible practical advantages over a simple full-parameter fine-tuning baseline. To explain this surprising effectiveness, we provide an in-depth mechanistic analysis showing that full fine-tuning succeeds via distributed micro-adjustments and more flexible reshaping of high-level semantic presentations without suffering from overfitting. Additionally, we quantify the performance collapse of multimodal models in specialized domains as a result of linguistic rarity using adjusted Zipf frequency scores. By releasing FEWTRANS, we aim to provide a rigorous "ruler" to streamline reproducible advances in few-shot transfer learning research. We make the FEWTRANS benchmark publicly available at https://github.com/Frankluox/FewTrans.
Problem

Research questions and friction points this paper is trying to address.

few-shot transfer
evaluation protocol
pre-trained models
benchmark
transferability
Innovation

Methods, ideas, or system contributions that make the work stand out.

few-shot transfer learning
evaluation protocol
hyperparameter ensemble
full fine-tuning
multimodal models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Xu Luo
Xu Luo
UESTC
Machine LearningRobotics
J
Ji Zhang
Southwest Jiaotong University, Chengdu, China
Lianli Gao
Lianli Gao
UESTC
Vision and Language
H
Heng Tao Shen
Tongji University, Shanghai, China
J
Jingkuan Song
Tongji University, Shanghai, China