meta-transfer learning

Designs and evaluates meta-learned parameter initializations and associated adaptation procedures that transfer across tasks so models can be rapidly fine-tuned with few steps; builds episodic meta-training pipelines, warm-start strategies, and optimization schedules that maximize early per-task performance and sustained adaptation while minimizing meta-training budget and cost.

meta-transferlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters

Feb 04, 2024
AS
Arsalan Sharifnassab
🏛️ University of Alberta | Leiden University

To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.

Dynamically adjusting step sizes during model optimizationOptimizing meta-parameters for efficient machine learning trainingReducing regret by considering long-term impact of learning rates

Meta-Learning Adaptable Foundation Models

Oct 29, 2024
JL
Jacob L. Block
🏛️ The University of Texas at Austin

Standard fine-tuning of foundation models suffers from low downstream adaptation efficiency and fails to recover the optimal adaptable parameter set. Method: We propose the first PEFT co-optimization framework that explicitly integrates meta-learning (MAML-style) into the foundation model’s retraining phase, using a LoRA-inspired low-rank adaptation structure. Contribution/Results: We theoretically prove that standard retraining is inherently suboptimal in adaptability, whereas our method strictly recovers the optimal adaptable parameters and provides a generalization error bound. Experiments on RoBERTa with the ConvAI2 dialogue continuation task demonstrate significant improvements in zero-shot and few-shot rapid adaptation performance, empirically validating the theoretical guidance.

Analyzing theoretical guarantees for parameter-efficient fine-tuning on unseen tasksDemonstrating performance improvements over standard retraining methods empiricallyDeveloping provable meta-learning for low-rank adaptation of foundation models

Optimizing ML Training with Metagradient Descent

Mar 17, 2025
LE
Logan Engstrom
🏛️ MIT | Stanford | UIUC

To address the challenge of efficiently optimizing high-dimensional configuration spaces in large-scale machine learning training, this paper proposes a scalable meta-gradient computation algorithm and the Smooth Model Training (SMT) framework—enabling, for the first time, end-to-end, differentiable joint optimization of training strategies. Methodologically, it integrates reverse-mode automatic differentiation through training loops, smooth modeling of training trajectories, and meta-gradient descent (MGD) to jointly optimize data selection, poisoning-resilient strategies, and learning rate scheduling. Key contributions are: (1) a breakthrough in scalable meta-gradient computation for large-scale training; and (2) the SMT framework, which ensures stability and convergence of MGD under realistic dynamic training conditions. Experiments demonstrate that the proposed data selection method significantly outperforms existing approaches; robustness against accuracy-degrading data poisoning attacks improves by an order of magnitude; and the fully automated learning rate scheduler matches or exceeds hand-crafted designs in performance.

Efficiently calculating metagradients for model trainingImproving dataset selection and learning rate schedulesOptimizing training setup for large-scale ML models

In large-scale pretraining, learning rate scheduling critically influences both training efficiency and model performance. This work proposes two paradigms—Fitting and Transfer. The Fitting paradigm establishes, for the first time, a scaling law for learning rate search factors, reducing hyperparameter tuning complexity from O(n³) to O(n·C_D·C_η). The Transfer paradigm extends μTransfer to Mixture-of-Experts (MoE) architectures and generalizes it across multiple hyperparameter dimensions, including depth, weight decay, and token length. Empirical results demonstrate that while μTransfer exhibits limited scalability in large-scale settings, the Fitting paradigm—grounded in the derived scaling law—offers superior scalability and practicality, providing a systematic guideline for hyperparameter tuning in industrial-scale pretraining.

hyperparameter optimizationlarge-scale pre-traininglearning rate

Rethinking Meta-Learning from a Learning Lens

Sep 13, 2024
JW
Jingyao Wang
🏛️ University of Chinese Academy of Sciences | Institute of Software Chinese Academy of Sciences

Meta-learning models often suffer from overfitting to training tasks and poor generalization, stemming from task-wise co-adaptation that induces dual risks—both overfitting and underfitting. Method: This work systematically analyzes error sources from a learning dynamics perspective and proposes a task-relation-driven calibration paradigm: (i) constructing a task relationship matrix; (ii) designing relation-aware consistency regularization; (iii) introducing meta-data-driven task similarity estimation; and (iv) conducting theory-guided optimization stability analysis. Based on this, we develop TRLearner—a plug-and-play method requiring no architectural or data modifications. Contribution/Results: TRLearner significantly improves generalization across multiple benchmarks. Theoretically, it ensures enhanced convergence guarantees; empirically, stronger task similarity yields more pronounced collaborative gains, validating the efficacy of relation-aware calibration.

Meta-learning overfits on training tasks.Task relations optimize meta-learning process.Task similarity affects model adaptation effectiveness.

Latest Papers

What's happening recently
View more

This work addresses the challenges of catastrophic forgetting and task-specific knowledge dilution in continual fine-tuning of large language models. Existing approaches typically rely on experience replay or task-specific adapters, incurring substantial computational and storage overhead. To overcome these limitations, the authors propose a novel paradigm that requires neither replay nor additional adapter modules. Their method employs a brief warm-up fine-tuning phase, followed by identification of a core subset of parameters per task using parameter importance metrics—such as L2 norm and Fisher information—and task-specificity analysis based on cosine similarity of update directions. During subsequent training, only this critical parameter subset is updated while the rest remain frozen to preserve prior knowledge. Extensive experiments demonstrate that this approach significantly outperforms current state-of-the-art methods across multiple benchmarks, confirming its effectiveness for large-scale models under resource constraints and its transferability across different model sizes.

catastrophic forgettingcontinual fine-tuninglarge language models

This work addresses the slow adaptation and limited generalization of conventional reinforcement learning in multi-task and non-stationary energy systems by proposing a novel meta-reinforcement learning framework. The approach integrates bilevel optimization with a hybrid Actor-Critic architecture, jointly optimizing a shared state feature extractor and incorporating a parameter-sharing mechanism between inner- and outer-loop policy networks to significantly enhance sample efficiency and cross-task adaptability. Experimental evaluation on a decade-long real-world building energy management dataset demonstrates that the proposed framework achieves faster adaptation upon task revisitation and superior control performance compared to existing reinforcement learning and meta-reinforcement learning methods.

Energy SystemsFast AdaptationMeta-Reinforcement Learning

This study addresses the difficulty language model agents face in efficiently adapting execution frameworks to diverse tasks at test time. To this end, this work proposes "framework learning," which formulates framework revision as meta-learning over executable programs. Specifically, a proposer model is trained via reinforcement learning to iteratively refine a solver's code framework using execution feedback, thereby enabling test-time adaptation without parameter updates. By integrating large language model agents with program synthesis and automated repair techniques, this approach endows agents with the capacity to continuously generalize and improve from experience. Experimental results demonstrate significant performance gains on reasoning and multi-hop question answering tasks, validating that such test-time adaptation capabilities transfer effectively to unseen tasks.

Executable ProgramsHarness LearningMeta-Learning

This study addresses the challenge that a single program struggles to accommodate heterogeneous requests and that manual partitioning is inefficient. To this end, we propose an adaptive prompt optimization framework that, for the first time, unifies request routing and program evolution within a shared search budget by jointly evolving a router and an expert program library. By integrating execution-trace-based reflection, genetic programming, and natural language instruction editing, the method leverages human-readable feedback to achieve label-free automatic alignment and knowledge inheritance. Evaluated on Qwen3-8B, the proposed framework attains 100% routing accuracy and improves the family-average test score from 52.6 to 70.6, significantly outperforming baseline methods such as GEPA and GRPO.

adaptive routingheterogeneous requestsprogram synthesis

This work addresses the challenge of reference trajectory tracking for uncertain nonlinear systems with limited data by proposing a meta-learning-based control framework. The approach learns a shared dynamic representation from structurally similar source systems during an offline phase and enables rapid adaptation of the controller to a new target system using only a few online data samples. Innovatively adapting implicit Model-Agnostic Meta-Learning (iMAML) to the control domain, the method establishes a general bilevel optimization framework compatible with diverse learning algorithms while significantly reducing memory overhead and approximation error. Two implementation pathways—neural state-space models and deep Q-networks, corresponding respectively to explicit and implicit system identification—are evaluated through simulations and hardware experiments, consistently demonstrating superior control performance over baseline methods and confirming the framework’s effectiveness and practicality.

data-efficient controlmeta-learningreference tracking

Hot Scholars

HT

Hao Tan

Adobe Research
Vision and Language3D Multimodal
TQ

Tieyun Qian

Wuhan University
natural language processingweb data mining
JW

Jinlin Wang

DeepWisdom
Computer Vision、Multi-Agent System、Large Language Model、Large Vision-Language Model