adaptive instance weighting

Design and implement methods that compute per-instance scores or weights and apply them to training, guidance, or evaluation processes to prioritize, downweight, or emphasize particular examples. This includes algorithms for instance-level weighting, adaptive guidance weighting, multiple-instance learning and aggregation, constructing or scoring hard/worst-case instances, and using techniques such as GMM-based weighting or reliability scoring to balance imitation and autonomous exploration.

adaptiveinstanceweighting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of dynamically adapting personalized exercise recommendations to learners’ evolving skill levels in large-scale online education. The authors propose a contextual Thompson sampling-based multi-armed bandit approach that integrates learner characteristics and historical performance to perform real-time Bayesian inference of skill proficiency. By optimizing exercise sequences with the explicit objective of maximizing skill gain, the method tailors recommendations to individual learning trajectories. Experiments on real-world data from a mathematics tutoring platform demonstrate that the proposed approach significantly enhances skill acquisition, effectively accommodates individual differences, and simultaneously identifies high-impact exercises and at-risk learners requiring intervention. The solution exhibits strong scalability and practical utility for real-world educational applications.

adaptive practiceeducational recommender systemslearner modeling

Leveraging weights signals - Predicting and improving generalizability in reinforcement learning

Nov 25, 2025
OM
Olivier Moulin
🏛️ Vrije Universiteit Amsterdam | Vrije Universiteit Medical Center Amsterdam

Reinforcement learning agents often exhibit poor generalization to unseen environments due to overfitting to training conditions. To address this, we propose the first method that predicts agent generalization performance directly from neural network weight signals and integrates this prediction into the PPO objective for generalization-aware policy optimization. Our key contributions are: (1) a differentiable weight-feature extraction module that maps model parameters to a scalar generalization score; and (2) a generalization-aware regularization term incorporated into the PPO loss, which explicitly encourages learning of robust, environment-invariant representations. Experiments across diverse generalization benchmarks—including visual-observation domains (ProcGen) and dynamics-shift settings (MultiRoom)—demonstrate substantial improvements in cross-environment performance: our method achieves an average generalization score 23.6% higher than standard PPO, without requiring environmental augmentation, domain randomization, or auxiliary supervision.

Addressing overfitting in reinforcement learning training environmentsImproving generalization by modifying PPO loss functionPredicting RL agent generalizability using neural network weight signals

MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters

Feb 04, 2024
AS
Arsalan Sharifnassab
🏛️ University of Alberta | Leiden University

To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.

Dynamically adjusting step sizes during model optimizationOptimizing meta-parameters for efficient machine learning trainingReducing regret by considering long-term impact of learning rates

Instance-Conditioned Adaptation for Large-scale Generalization of Neural Combinatorial Optimization

May 03, 2024
CZ
Changliang Zhou
🏛️ Southern University of Science and Technology | City University of Hong Kong | Huawei

Existing neural combinatorial optimization (NCO) methods exhibit poor generalization to large-scale routing problems—such as the Traveling Salesman Problem (TSP) and Capacitated Vehicle Routing Problem (CVRP)—limiting their applicability in real-world intelligent transportation systems. To address this, we propose Instance-Conditional Adaptive Mechanism (ICAM), a construction-based graph neural network model that achieves cross-scale adaptability via lightweight adapters conditioned on instance-specific embeddings. We further introduce a novel three-stage unsupervised reinforcement learning paradigm, enabling end-to-end training on instances ranging from 100 to 1,000 nodes without access to optimal solution labels. Experiments demonstrate that ICAM achieves state-of-the-art performance among construction-based NCO approaches on TSP and CVRP benchmarks, scales robustly up to 1,000 nodes, and delivers highly efficient inference—significantly outperforming existing methods.

Enhancing solution quality across different problem scalesImproving large-scale generalization of neural routing solversReducing time and memory overhead in NCO methods

Optimization of Scoring Rules

Jul 06, 2020
YL
Yingkai Li
🏛️ Yale University | Northwestern University | Toyota Technological Institute at Chicago

This paper addresses the design of proper scoring rules for multidimensional forecasting settings, aiming to incentivize forecasters to exert effort and truthfully report their beliefs. Methodologically, it introduces the first optimization framework explicitly targeting *effort incentives*, integrating game-theoretic modeling with convex optimization. For simple settings, it derives closed-form characterizations of optimal rules; for general cases, it develops an efficient and exact algorithm; and it identifies several structurally simple approximate rules with near-optimal performance. Theoretical analysis reveals that classical proper scoring rules—such as the quadratic score—can substantially deviate from optimality under multidimensional effort. In contrast, the proposed algorithm computes exact optimal rules, while the simple approximations achieve over 95% of the optimal incentive efficiency. These results establish a new paradigm for information design and prediction market mechanisms, bridging incentive alignment with practical implementability.

Comparing optimal scoring rules with standard alternativesDesigning incentives for multi-dimensional information acquisitionOptimizing scoring rules for truthful information reporting

Latest Papers

What's happening recently
View more

This work addresses the limitation of conventional metadata-based curriculum learning in identifying training scenarios critical for improving motion planning performance in interactive driving tasks. For the first time, it introduces gradient-driven data valuation (TracIn) to this domain, constructing a curriculum by quantifying each training sample’s contribution to validation loss reduction. The resulting curriculum is shown to be nearly orthogonal to handcrafted metadata, capturing dynamic training effects that metadata overlooks. Evaluated on the GameFormer architecture using the nuPlan benchmark, the proposed approach achieves an average ADE of 1.704 ± 0.029 meters under curriculum-weighted training, significantly outperforming metadata-based curriculum learning (1.822 ± 0.014 meters, p = 0.021) and exhibiting lower variance than uniform sampling.

curriculum learningdata valuationgame-theoretic motion planning

This work addresses the limited generalization and low adaptation efficiency of initial policy priors in meta-reinforcement learning by introducing quasi-Monte Carlo (QMC) sampling into the weight initialization phase of meta-learning for the first time. By integrating bounded population search with top-prior aggregation, the proposed approach constructs an effective initialization strategy for continuous control tasks. Experimental results demonstrate that QMC-based initialization significantly accelerates training convergence in unseen environments with high task similarity, whereas conventional orthogonal initialization retains an advantage when tasks are substantially dissimilar. This study reveals the critical influence of task similarity on the effectiveness of initialization strategies and offers a novel perspective for enabling efficient warm-starting in meta-reinforcement learning.

continuous controlmeta-reinforcement learningquasi-Monte Carlo

This study addresses a critical limitation in existing experimental designs for evaluating algorithm-assisted decision-making: the neglect of human behavioral adaptations—such as automation bias and alert fatigue—that arise from repeated exposure to algorithmic recommendations, leading to biased effect estimates. To remedy this, the paper introduces, for the first time, a systematic causal estimand tailored to repeated-exposure settings and proposes a minimax staircase double-wedge randomized experimental framework. This approach enables unbiased estimation of the true effect of algorithmic assistance, whereas conventional designs exhibit substantial bias under typical adaptation patterns. The proposed method thus demonstrates superior accuracy and robustness in capturing the genuine impact of algorithmic interventions in dynamic human–AI interaction contexts.

algorithm-assisted decision-makingbehavioral adaptationeffect estimands

This work addresses the limited generalization capability of large language models (LLMs) in automatically composing algorithms under few-shot settings by proposing a Potential-aware Instance and Algorithm Co-evolution framework (PIAC). PIAC introduces a potential gain metric that evaluates instance difficulty without requiring ground-truth solutions and leverages LLMs to generate diverse instance variation operators, thereby overcoming the reliance on high-quality solutions and unimodal generation inherent in prior approaches. The framework flexibly integrates multiple algorithmic components—including greedy construction, ant colony optimization, and guided local search—and demonstrates significant performance gains over existing LLM-based algorithm composition methods across six datasets with varying distributions for the Traveling Salesman Problem (TSP) and Capacitated Vehicle Routing Problem (CVRP). Notably, the greedy construction variant achieves a relative improvement of 19.76% on TSP instances.

algorithm portfoliocombinatorial optimizationgeneralization

This work addresses the limitations of existing self-training methods for fine-tuning large language models, which are highly sensitive to synthetic data quality and suffer from diminishing margins between positive and negative samples during iterative optimization. To overcome these challenges, the authors propose the TPAW algorithm, which operates in a fully self-supervised setting by constructing a cooperative-competitive ensemble composed of the current policy model and historical checkpoints to engage in self-play. TPAW incorporates a dual adaptive weighting mechanism—comprising response reweighting and participant dynamic weighting—to enhance training stability and alignment efficacy. Requiring no human supervision and initialized solely from a supervised fine-tuned (SFT) model, TPAW iteratively refines model performance and consistently outperforms state-of-the-art baselines across multiple base models and LLM benchmarks, significantly improving alignment outcomes.

bias amplificationLLM alignmentoptimization gap

Hot Scholars

MV

Maria Vakalopoulou

Assistant Professor at CentraleSupélec
Medical ImagingRemote SensingComputer VisionMachine Learning
MV

Michal Valko

Chief Models Officer @ Stealth Startup, Inria & MVA - Ex: Llama at Meta; Gemini and BYOL @ Deepmind
large language modelsreasoningfine-tuningtest-time computation
NN

Nassir Navab

Professor of Computer Science, Technische Universität München
CM

Carsten Marr

Institute of AI for Health @ Helmholtz Munich & Clinics @ LMU München
AI for Biomed & Health