Score
Designs and trains meta-weighting models or meta-learners that output per-example or per-batch weights and the bi-level optimization (meta-training) procedures that learn those weights from a validation objective. Builds and analyzes the end-to-end training loop, meta-gradient computation, and data-weight update/normalization rules to select or weight training examples and thereby influence optimization dynamics and final validation performance.
To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.
Existing data-free meta-learning methods are constrained to parameter-space optimization and require homogeneous model architectures, limiting scalability to large-scale pretrained models. This paper introduces the first data-free meta-learning framework tailored for heterogeneous pretrained models, enabling extraction of implicit prior knowledge without access to original training data. Our approach features two core innovations: Episode Curriculum Inversion (ECI) and Inversion Calibration Following Inner Loop (ICFIL). Leveraging pseudo-task distillation, adversarial end-to-end meta-training, and curriculum-based pseudo-episode generation, the framework achieves generalizable meta-adaptation across architectural heterogeneity, model scales (up to 10B parameters), and diverse datasets. Experiments demonstrate substantial improvements over state-of-the-art data-free meta-learning methods across multiple benchmarks, validating both strong generalization and seamless scalability.
Meta-learning models often suffer from overfitting to training tasks and poor generalization, stemming from task-wise co-adaptation that induces dual risks—both overfitting and underfitting. Method: This work systematically analyzes error sources from a learning dynamics perspective and proposes a task-relation-driven calibration paradigm: (i) constructing a task relationship matrix; (ii) designing relation-aware consistency regularization; (iii) introducing meta-data-driven task similarity estimation; and (iv) conducting theory-guided optimization stability analysis. Based on this, we develop TRLearner—a plug-and-play method requiring no architectural or data modifications. Contribution/Results: TRLearner significantly improves generalization across multiple benchmarks. Theoretically, it ensures enhanced convergence guarantees; empirically, stronger task similarity yields more pronounced collaborative gains, validating the efficacy of relation-aware calibration.
In meta-learning, manually constructing validation sets suffers from poor scalability with increasing classes, difficulty in simultaneously ensuring class balance and label reliability, and dependence on human curation. To address these issues, this paper proposes a utility-driven validation set construction paradigm. We formally define validation set utility along three dimensions: informativeness, necessity (i.e., class balance), and label reliability—introducing the first end-to-end differentiable algorithm, INOLML, optimized for this utility metric. INOLML jointly performs utility-aware validation set selection, self-supervised label correction, and imbalance- and noise-robust meta-training. Extensive experiments on multiple benchmark datasets demonstrate substantial improvements over state-of-the-art methods, establishing new SOTA performance for meta-learning under imbalanced and noisy labels.
To bridge the gap between data-hungry deep models and human-like few-shot learning efficiency, this paper investigates the long-overlooked loss function component within meta-learning frameworks and proposes a novel dynamic adaptive loss learning paradigm. Methodologically, it introduces (1) EvoMAL—an interpretable symbolic loss evolution method that integrates symbolic regression with evolutionary algorithms to generate lightweight, task-adaptive, and interpretable loss functions; (2) Sparse Label Smoothing Regularization (SparseLSR), a new regularization technique for mitigating label noise in low-data regimes; and (3) NPBML—a unified framework jointly optimizing meta-initialization, meta-optimizer, and loss function. Experiments across multiple few-shot benchmarks demonstrate state-of-the-art performance: classification accuracy improves significantly, while memory overhead for loss learning is reduced by over 80%.
This work addresses the limitations of existing meta-learning training data selection (MTS) methods, which suffer from performance degradation due to the mismatch between synthetic and real data distributions. For the first time, it identifies the underlying failure mechanisms of MTS through the lenses of low gradient signal-to-noise ratio and insufficient feature informativeness. To overcome these issues without requiring complex architectural modifications, the study proposes an efficient strategy that enhances gradient quality by increasing batch size and introduces a novel data quality metric to improve selection efficacy. Built upon a bilevel optimization framework, the method jointly models data distributional positioning and training dynamics. Experiments across four benchmarks demonstrate average performance gains of 5.49% over non-selection baselines and 2.89% over the strongest existing MTS approach, substantially improving the utilization efficiency of synthetic data.
Existing meta-black-box optimization (MetaBBO) methods rely on fixed benchmark suites (e.g., BBOB) for training, leading to overfitting and limited generalization due to insufficient problem diversity. To address this, we propose LSRE—a latent-space reverse-engineering framework. First, an autoencoder compresses problem features into a two-dimensional latent space. Second, uniform grid sampling in this space, combined with genetic programming-based inversion, synthesizes a highly diverse problem suite, Diverse-BBO—the first approach to construct optimization problems via latent-space inverse engineering. An L2-distance constraint ensures both executability and distributional fidelity of generated problems. Experiments demonstrate that MetaBBO models trained on Diverse-BBO achieve significant performance gains over baselines on both synthetic and real-world tasks. Ablation studies confirm that enhanced problem diversity is critical for improving generalization.
This work addresses the longstanding challenge in optimization of simultaneously achieving stability and scalability in high-dimensional, long-horizon training settings: gradient-based methods often suffer from instability, while gradient-free approaches struggle to scale to large parameter spaces. To overcome this, the paper proposes a bilevel meta-learning algorithm wherein an inner loop performs adaptive updates in the high-dimensional parameter space, guided efficiently by an outer loop that optimizes a small set of meta-parameters via zeroth-order methods. This architecture effectively decouples the complexity of high-dimensional learning from the control of learning dynamics, thereby circumventing the dual limitations of temporal horizon and dimensionality inherent in existing approaches. As a result, the method substantially enhances the stability, efficiency, and robustness of large-scale models during prolonged training.
This work addresses the challenge of reference trajectory tracking for uncertain nonlinear systems with limited data by proposing a meta-learning-based control framework. The approach learns a shared dynamic representation from structurally similar source systems during an offline phase and enables rapid adaptation of the controller to a new target system using only a few online data samples. Innovatively adapting implicit Model-Agnostic Meta-Learning (iMAML) to the control domain, the method establishes a general bilevel optimization framework compatible with diverse learning algorithms while significantly reducing memory overhead and approximation error. Two implementation pathways—neural state-space models and deep Q-networks, corresponding respectively to explicit and implicit system identification—are evaluated through simulations and hardware experiments, consistently demonstrating superior control performance over baseline methods and confirming the framework’s effectiveness and practicality.
This work addresses the challenge that large language models struggle to adapt and self-improve on specific tasks during inference. To this end, the authors propose the MASS framework, which introduces the first end-to-end test-time meta-self-adaptation mechanism. MASS employs a bilevel optimization architecture: in the inner loop, the model is updated using self-generated, task-specific synthetic data, while the outer loop meta-learns an optimal data generation strategy and reward signal. This approach substantially enhances the model’s test-time adaptability and data efficiency on tasks such as mathematical reasoning, enabling the effective generation of instance-level curriculum data and yielding significant performance gains.