Score
Designs and evaluates meta-learned parameter initializations and associated adaptation procedures that transfer across tasks so models can be rapidly fine-tuned with few steps; builds episodic meta-training pipelines, warm-start strategies, and optimization schedules that maximize early per-task performance and sustained adaptation while minimizing meta-training budget and cost.
To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.
Standard fine-tuning of foundation models suffers from low downstream adaptation efficiency and fails to recover the optimal adaptable parameter set. Method: We propose the first PEFT co-optimization framework that explicitly integrates meta-learning (MAML-style) into the foundation model’s retraining phase, using a LoRA-inspired low-rank adaptation structure. Contribution/Results: We theoretically prove that standard retraining is inherently suboptimal in adaptability, whereas our method strictly recovers the optimal adaptable parameters and provides a generalization error bound. Experiments on RoBERTa with the ConvAI2 dialogue continuation task demonstrate significant improvements in zero-shot and few-shot rapid adaptation performance, empirically validating the theoretical guidance.
To address the challenge of efficiently optimizing high-dimensional configuration spaces in large-scale machine learning training, this paper proposes a scalable meta-gradient computation algorithm and the Smooth Model Training (SMT) framework—enabling, for the first time, end-to-end, differentiable joint optimization of training strategies. Methodologically, it integrates reverse-mode automatic differentiation through training loops, smooth modeling of training trajectories, and meta-gradient descent (MGD) to jointly optimize data selection, poisoning-resilient strategies, and learning rate scheduling. Key contributions are: (1) a breakthrough in scalable meta-gradient computation for large-scale training; and (2) the SMT framework, which ensures stability and convergence of MGD under realistic dynamic training conditions. Experiments demonstrate that the proposed data selection method significantly outperforms existing approaches; robustness against accuracy-degrading data poisoning attacks improves by an order of magnitude; and the fully automated learning rate scheduler matches or exceeds hand-crafted designs in performance.
In large-scale pretraining, learning rate scheduling critically influences both training efficiency and model performance. This work proposes two paradigms—Fitting and Transfer. The Fitting paradigm establishes, for the first time, a scaling law for learning rate search factors, reducing hyperparameter tuning complexity from O(n³) to O(n·C_D·C_η). The Transfer paradigm extends μTransfer to Mixture-of-Experts (MoE) architectures and generalizes it across multiple hyperparameter dimensions, including depth, weight decay, and token length. Empirical results demonstrate that while μTransfer exhibits limited scalability in large-scale settings, the Fitting paradigm—grounded in the derived scaling law—offers superior scalability and practicality, providing a systematic guideline for hyperparameter tuning in industrial-scale pretraining.
Meta-learning models often suffer from overfitting to training tasks and poor generalization, stemming from task-wise co-adaptation that induces dual risks—both overfitting and underfitting. Method: This work systematically analyzes error sources from a learning dynamics perspective and proposes a task-relation-driven calibration paradigm: (i) constructing a task relationship matrix; (ii) designing relation-aware consistency regularization; (iii) introducing meta-data-driven task similarity estimation; and (iv) conducting theory-guided optimization stability analysis. Based on this, we develop TRLearner—a plug-and-play method requiring no architectural or data modifications. Contribution/Results: TRLearner significantly improves generalization across multiple benchmarks. Theoretically, it ensures enhanced convergence guarantees; empirically, stronger task similarity yields more pronounced collaborative gains, validating the efficacy of relation-aware calibration.
This work addresses the challenges of catastrophic forgetting and task-specific knowledge dilution in continual fine-tuning of large language models. Existing approaches typically rely on experience replay or task-specific adapters, incurring substantial computational and storage overhead. To overcome these limitations, the authors propose a novel paradigm that requires neither replay nor additional adapter modules. Their method employs a brief warm-up fine-tuning phase, followed by identification of a core subset of parameters per task using parameter importance metrics—such as L2 norm and Fisher information—and task-specificity analysis based on cosine similarity of update directions. During subsequent training, only this critical parameter subset is updated while the rest remain frozen to preserve prior knowledge. Extensive experiments demonstrate that this approach significantly outperforms current state-of-the-art methods across multiple benchmarks, confirming its effectiveness for large-scale models under resource constraints and its transferability across different model sizes.
This work addresses the slow adaptation and limited generalization of conventional reinforcement learning in multi-task and non-stationary energy systems by proposing a novel meta-reinforcement learning framework. The approach integrates bilevel optimization with a hybrid Actor-Critic architecture, jointly optimizing a shared state feature extractor and incorporating a parameter-sharing mechanism between inner- and outer-loop policy networks to significantly enhance sample efficiency and cross-task adaptability. Experimental evaluation on a decade-long real-world building energy management dataset demonstrates that the proposed framework achieves faster adaptation upon task revisitation and superior control performance compared to existing reinforcement learning and meta-reinforcement learning methods.
This study addresses the difficulty language model agents face in efficiently adapting execution frameworks to diverse tasks at test time. To this end, this work proposes "framework learning," which formulates framework revision as meta-learning over executable programs. Specifically, a proposer model is trained via reinforcement learning to iteratively refine a solver's code framework using execution feedback, thereby enabling test-time adaptation without parameter updates. By integrating large language model agents with program synthesis and automated repair techniques, this approach endows agents with the capacity to continuously generalize and improve from experience. Experimental results demonstrate significant performance gains on reasoning and multi-hop question answering tasks, validating that such test-time adaptation capabilities transfer effectively to unseen tasks.
This study addresses the challenge that a single program struggles to accommodate heterogeneous requests and that manual partitioning is inefficient. To this end, we propose an adaptive prompt optimization framework that, for the first time, unifies request routing and program evolution within a shared search budget by jointly evolving a router and an expert program library. By integrating execution-trace-based reflection, genetic programming, and natural language instruction editing, the method leverages human-readable feedback to achieve label-free automatic alignment and knowledge inheritance. Evaluated on Qwen3-8B, the proposed framework attains 100% routing accuracy and improves the family-average test score from 52.6 to 70.6, significantly outperforming baseline methods such as GEPA and GRPO.
This work addresses the challenge of reference trajectory tracking for uncertain nonlinear systems with limited data by proposing a meta-learning-based control framework. The approach learns a shared dynamic representation from structurally similar source systems during an offline phase and enables rapid adaptation of the controller to a new target system using only a few online data samples. Innovatively adapting implicit Model-Agnostic Meta-Learning (iMAML) to the control domain, the method establishes a general bilevel optimization framework compatible with diverse learning algorithms while significantly reducing memory overhead and approximation error. Two implementation pathways—neural state-space models and deep Q-networks, corresponding respectively to explicit and implicit system identification—are evaluated through simulations and hardware experiments, consistently demonstrating superior control performance over baseline methods and confirming the framework’s effectiveness and practicality.