Understanding Optimal Feature Transfer via a Fine-Grained Bias-Variance Analysis

📅 2024-04-18
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates the optimality of pretrained feature representations for downstream few-shot learning tasks in transfer learning. Methodologically, it establishes a linear feature transfer model and derives an asymptotic bias–variance decomposition of the downstream risk. Theoretically, it is the first to demonstrate that, on average, the optimal pretrained representation intrinsically exhibits sparsity, and that a phase transition emerges—from hard-thresholding feature selection to soft weighting—without explicit sparse regularization. The analysis integrates asymptotic statistics, multi-task averaging optimization, and linear transfer modeling. The theory precisely characterizes the underlying mechanisms driving both sparsity and the phase transition. Empirical validation on image and text few-shot benchmarks confirms substantial generalization gains over strong baselines. Collectively, this work provides a novel theoretical lens for understanding the intrinsic effectiveness of pretrained representations.

Technology Category

Machine Learning: Transfer, Domain Adaptation, Multi-Task LearningComputer Vision: Representation Learning for VisionNatural Language Processing: Learning & Optimization for NLP

Application Category

Web Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphs
📝 Abstract
In the transfer learning paradigm models learn useful representations (or features) during a data-rich pretraining stage, and then use the pretrained representation to improve model performance on data-scarce downstream tasks. In this work, we explore transfer learning with the goal of optimizing downstream performance. We introduce a simple linear model that takes as input an arbitrary pretrained feature transform. We derive exact asymptotics of the downstream risk and its extit{fine-grained} bias-variance decomposition. We then identify the pretrained representation that optimizes the asymptotic downstream bias and variance averaged over an ensemble of downstream tasks. Our theoretical and empirical analysis uncovers the surprising phenomenon that the optimal featurization is naturally sparse, even in the absence of explicit sparsity-inducing priors or penalties. Additionally, we identify a phase transition where the optimal pretrained representation shifts from hard selection to soft selection of relevant features.
Problem

Research questions and friction points this paper is trying to address.

Optimizing downstream performance in transfer learning
Analyzing bias-variance decomposition of downstream risk
Identifying optimal sparse featurization for downstream tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Linear model with pretrained feature transform
Exact asymptotics of downstream risk
Optimal sparse featurization without sparsity priors
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Harvard University | Google DeepMind
Y
Yufan Li
Department of Statistics, Harvard University
S
Subhabrata Sen
Department of Statistics, Harvard University
Ben Adlam
Ben Adlam
Google DeepMind