dataset optimization

Designs and implements algorithms and procedures to select, weight, and schedule training examples or scenarios so as to improve a model’s generalization and learning efficiency. This includes building gradient-based data‑selection and loss‑sensitivity analyses, tuning training schedules, and iteratively updating scenario sets or dataset composition.

datasetoptimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Optimizing ML Training with Metagradient Descent

Mar 17, 2025
LE
Logan Engstrom
🏛️ MIT | Stanford | UIUC

To address the challenge of efficiently optimizing high-dimensional configuration spaces in large-scale machine learning training, this paper proposes a scalable meta-gradient computation algorithm and the Smooth Model Training (SMT) framework—enabling, for the first time, end-to-end, differentiable joint optimization of training strategies. Methodologically, it integrates reverse-mode automatic differentiation through training loops, smooth modeling of training trajectories, and meta-gradient descent (MGD) to jointly optimize data selection, poisoning-resilient strategies, and learning rate scheduling. Key contributions are: (1) a breakthrough in scalable meta-gradient computation for large-scale training; and (2) the SMT framework, which ensures stability and convergence of MGD under realistic dynamic training conditions. Experiments demonstrate that the proposed data selection method significantly outperforms existing approaches; robustness against accuracy-degrading data poisoning attacks improves by an order of magnitude; and the fully automated learning rate scheduler matches or exceeds hand-crafted designs in performance.

Efficiently calculating metagradients for model trainingImproving dataset selection and learning rate schedulesOptimizing training setup for large-scale ML models

Step-Opt: Boosting Optimization Modeling in LLMs through Iterative Data Synthesis and Structured Validation

Jun 21, 2025
YW
Yang Wu
🏛️ C2DL | Institute of Automation | Chinese Academy of Sciences | School of Artificial Intelligence | University of Chinese Academy of Sciences | Dalian Minzu University

Large language models (LLMs) struggle with complex problem comprehension and precise mathematical modeling in operations research (OR) optimization tasks. To address this, we propose Step-Opt-Instruct—a novel framework that iteratively generates OR problems of progressively increasing complexity and incorporates a structured, stepwise validation mechanism to effectively prevent error propagation and enhance synthetic data quality. We apply supervised fine-tuning using this framework on LLaMA-3-8B and Mistral-7B. Experimental results demonstrate state-of-the-art performance across three major benchmarks—NL4OPT, MAMO, and IndustryOR—with a 17.01% improvement in micro-averaged accuracy on complex problems. The approach significantly strengthens generalization capability for multi-constraint, multi-objective decision-making tasks, advancing the frontier of natural-language-to-optimization modeling.

Enhancing LLMs for complex optimization modeling tasksGenerating high-quality data via iterative synthesis and validationImproving accuracy in Operations Research problem-solving

This paper challenges the unverified implicit assumption in the predict-then-optimize paradigm that “higher prediction accuracy necessarily yields better downstream decisions,” particularly in multiclass classification settings. Method: We propose a controllable, interpretable multiclass prediction simulation framework that explicitly models error types and distributions, enabling systematic analysis of how classification errors affect decision quality in constrained optimization. Contribution/Results: Experiments on job scheduling and other combinatorial optimization tasks reveal a nonlinear relationship between prediction error and decision performance: improving prediction accuracy does not guarantee improved solution quality—and can even degrade decisions when error patterns shift. Our findings question the conventional coupling logic between prediction and optimization, providing theoretical foundations and practical guidance for designing, evaluating, and calibrating classifiers specifically tailored to decision objectives.

Assessing Predict-Then-Optimize performance in machine scheduling problemsEvaluating how prediction error affects optimization solution qualitySimulating multiclass classifier predictions for experimental analysis

MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters

Feb 04, 2024
AS
Arsalan Sharifnassab
🏛️ University of Alberta | Leiden University

To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.

Dynamically adjusting step sizes during model optimizationOptimizing meta-parameters for efficient machine learning trainingReducing regret by considering long-term impact of learning rates

Existing algorithm selection models exhibit limited generalization capabilities in real-world optimization scenarios, struggling to maintain consistent performance across diverse domains. This work presents the first systematic evaluation of cross-domain generalization between synthetic benchmarks (BBOB, CEC) and practical applications—specifically robotic trajectory optimization and UAV path planning—using an algorithm selection framework grounded in problem features and historical performance data, complemented by a carefully designed cross-benchmark experimental protocol. The study uncovers the failure mechanisms and success boundaries of current approaches when deployed in realistic settings, thereby providing crucial empirical insights for developing more robust and universally applicable algorithm selection systems.

Algorithm SelectionBenchmarkingGeneralization

Latest Papers

What's happening recently
View more

This work proposes the first end-to-end automated artificial intelligence research framework capable of fully automating the development pipeline from algorithmic idea generation to executable machine learning classifiers. The approach integrates structured meta-prompt engineering with large language model–based code generation, augmented by an automated evaluation and iterative optimization mechanism. Experimental results on twenty standard datasets from the Infinity-Bench benchmark demonstrate that multiple novel classifiers autonomously generated by the framework significantly outperform baseline methods implemented in scikit-learn. This study thus achieves, for the first time, complete automation of the entire workflow—from initial algorithmic conception to deployable, runnable code—marking a significant step toward self-driving AI research systems.

AI automationautomate AI researchend-to-end framework

Conditional updates of neural network weights for increased out of training performance

Dec 03, 2025
JS
J. Saynisch-Wagner
🏛️ GFZ Helmholtz Centre for Geosciences

Neural networks often suffer significant performance degradation under distributional shifts—such as out-of-distribution generalization, spatiotemporal extrapolation, and cross-domain transfer—due to mismatches between training and deployment data distributions. To address this, we propose a conditional dynamic weight update framework. Its core innovation is a weight anomaly regression mechanism: sensitive weight change patterns induced by distribution shifts are identified via subset retraining; an interpretable regression predictor is then constructed to map input features to weight increments; finally, model parameters are conditionally extrapolated. The method integrates weight difference extraction, regression modeling, and extrapolation techniques, and is empirically validated on multi-source climate observation datasets. Across temporal, spatial, and cross-domain extrapolation tasks, it substantially improves prediction accuracy and robustness on out-of-distribution data, while preserving interpretability and practical applicability.

Address pattern and regime shifts in training versus application dataEnable temporal, spatial, and cross-domain extrapolation of neural networksEnhance neural network performance on out-of-distribution data

This work addresses the challenge of jointly optimizing data and model configurations in large language model training, a task rendered difficult by their high coupling. To this end, we propose JoBS, the first method to enable efficient joint optimization by integrating a scaling law–informed performance predictor into Bayesian optimization and leveraging multi-fidelity evaluation to substantially reduce the cost of full-scale training. JoBS not only yields an optimal budget allocation strategy but also consistently outperforms baselines that optimize only data, only model hyperparameters, or existing multi-fidelity Bayesian optimization approaches, achieving superior performance across diverse large language model tasks under identical computational budgets.

chicken-and-egg dilemmadata configurationjoint optimization

This work addresses a key limitation in existing data selection methods, which typically employ fixed selection ratios and overlook the impact of dynamically adjusting data volume on training efficiency and generalization. The authors propose PODS, a plug-and-play oscillating data scheduling framework that extends data selection from “what to select” to “how much to select.” By alternately applying low-ratio regularization phases and high-ratio recovery phases, PODS dynamically balances optimization fidelity and implicit regularization. The framework is lightweight, task-agnostic, and compatible with diverse static and dynamic selection strategies as well as model architectures. Experiments demonstrate that PODS reduces training costs by 50% while improving accuracy on ImageNet-1k and accelerates instruction fine-tuning of large language models by over 2× without any performance degradation.

data selectiondata-volume schedulingoptimization fidelity

Hot Scholars

ZT

Zhengzhong Tu

Texas A&M University, Google Research, University of Texas at Austin
Agentic AITrustworthy AIEmbodied AI
WX

Wayne Xin Zhao

Professor, Renmin University of China
Recommender SystemNatural Language ProcessingLarge Language Model
XH

Xuming Hu

Assistant Professor, HKUST(GZ) / HKUST
Natural Language ProcessingLarge Language Model
CL

Chaojian Li

Hong Kong University of Science and Technology
Efficient AIHardware / software codesign
CW

Cheng Wan

Georgia Institute of Technology