mixed-data training

Design and implement training procedures, curricula, and data pipelines that combine clean and noisy examples into a single training process, including sampling strategies, loss weighting, and augmentation so models are trained on mixed-quality data. Build experiments and analyses that measure how such combined clean–noisy training affects model robustness, regularization, and architecture-dependent performance.

mixed-datatraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Mixtera: A Data Plane for Foundation Model Training

Feb 27, 2025
MB
Maximilian Böther
🏛️ ETH Zurich

To address inefficiencies in manual dataset management and challenges in optimizing mixed-data sampling strategies and ordering amid rapidly expanding multi-source datasets for large model training, this paper proposes the first declarative, centralized, read-only data plane architecture. The architecture decouples data management from training frameworks and enables metadata-driven, cross-attribute (e.g., language, source) dynamic data mixing. It natively integrates state-of-the-art algorithms such as Adaptive Data Optimization (ADO) and supports real-time adjustment of mixing policies via model feedback. The system employs a training-framework-agnostic abstraction layer, a distributed sampling engine, and hardware-aware optimizations for GH200 superchip clusters. Empirical evaluation on 256 GPUs demonstrates zero throughput bottlenecks, substantial improvements in LLM and VLM performance, and a tenfold reduction in experimental iteration cost for mixture strategy tuning.

Automates data management for model trainingOptimizes data mixture and training sequenceScales efficiently across large-scale systems

Merge to Mix: Mixing Datasets via Model Merging

May 21, 2025
ZS
Zhixu Silvia Tao
🏛️ Princeton University | Fujitsu

Dataset mixing for large language model fine-tuning typically relies on labor-intensive trial-and-error and incurs substantial computational overhead. Method: This paper proposes a zero-shot dataset composition selection method that leverages model merging as a proxy evaluator—specifically, employing weighted averaging and task vector fusion to predict downstream performance of candidate dataset mixtures, while jointly optimizing mixture weights. Unlike conventional heuristic strategies requiring repeated full fine-tuning, our approach eliminates the need for any fine-tuning during evaluation. Contribution/Results: Experiments across multiple benchmarks demonstrate that our method significantly outperforms existing dataset selection techniques. It achieves comparable or improved final fine-tuned model performance while reducing computational cost by approximately 70% in GPU-hours, enabling efficient, scalable dataset composition design.

Accelerating dataset mixture selection for fine-tuning large modelsEnhancing model performance without full fine-tuning on each mixtureReducing reliance on heuristics and trial-and-error in dataset composition

Understanding and Mitigating the Bias in Sample Selection for Learning with Noisy Labels

Jan 24, 2024
QW
Qi Wei
🏛️ Nanyang Technological University | Zhejiang University

In noisy label learning, sample selection suffers from dual biases: data bias (imbalanced selection sets) and training bias (error accumulation). To address these issues, this paper proposes ITEM, a noise-tolerant expert model. ITEM is the first to jointly model and mitigate both biases in a unified framework. It introduces a lightweight multi-expert robust network architecture, integrated with a dual-weighted class-discriminative sampler and a hybrid mini-batch training strategy. Additionally, an error-robust optimization mechanism is incorporated to enhance generalization under label noise. Extensive experiments on multiple benchmark datasets with synthetic and real-world label noise demonstrate that ITEM consistently outperforms state-of-the-art methods, achieving average accuracy gains of 3–5% while reducing parameter count by over 20%. The source code is publicly available.

Addresses bias in sample selection for noisy labelsMitigates data and training bias in selection methodsProposes ITEM model for debiased learning and robust performance

Impact of Noisy Supervision in Foundation Model Learning

Mar 11, 2024
HC
Hao Chen
🏛️ Carnegie Mellon University | Microsoft Research Asia | Southern University of Science and Technology | RIKEN | the University of Tokyo

This work investigates how label noise in pretraining data affects the generalization of foundation models, revealing that while such noise may improve in-distribution (ID) performance, it inevitably degrades out-of-distribution (OOD) generalization—primarily by distorting the feature space geometry. To address this, the authors introduce “noisy-model tuning” as a novel paradigm and propose NMTune, a generic, parameter-efficient feature-space calibration method applicable to both white-box and black-box models. Extensive evaluation across synthetic and real-world noisy datasets—including ImageNet-1K, YFCC15M, and CC12M—covers diverse pretraining paradigms (fully supervised and vision-language contrastive), model architectures, downstream tasks, and tuning strategies. Results demonstrate that NMTune consistently mitigates noise-induced degradation, significantly improving OOD generalization across vision and language models—including proprietary API-based models—without dependence on model scale or task specificity.

Analyzes impact of label noise in pre-training datasets on model generalization.Explores noise influence across various datasets, models, and tuning methods.Proposes NMTune method to mitigate noise effects on downstream tasks.

Improved Regularization and Robustness for Fine-tuning in Neural Networks

Nov 08, 2021
DL
Dongyue Li
🏛️ Northeastern University

To address overfitting and poor noise robustness in large-model fine-tuning under few-shot settings, this paper proposes a novel fine-tuning framework that jointly enhances generalization and robustness. Methodologically, it introduces the first layer-wise L2-distance constraint regularization, integrated with confidence-guided self-label correction and dynamic reweighting. Theoretically, it is the first to systematically incorporate PAC-Bayes generalization bounds and noise stability analysis into fine-tuning design. Extensive experiments across seven image classification benchmarks demonstrate that the framework achieves an average accuracy improvement of 1.76%, gains +0.75% under few-shot conditions, and significantly outperforms baselines by +3.56% under label noise—validating both its effectiveness and robustness.

Addresses overfitting in fine-tuning large pre-trained modelsEnhances robustness against noisy labels in transfer learningProposes regularization techniques to improve generalization performance

Latest Papers

What's happening recently
View more

This work systematically investigates the impact of loss weighting strategies and output parameterizations on model performance in flow matching. Through numerical experiments on both synthetic data with controllable geometric structures and real-world images, the study disentangles their interaction effects across varying data manifold dimensions, model architectures, and dataset scales, using PSNR and FID as evaluation metrics. The analysis reveals, for the first time, how the optimal choice of loss weighting and parameterization depends critically on the intrinsic structure of the data. Building on these insights, the authors formulate practical design principles that substantially improve denoising accuracy and generation quality.

denoisingflow matchinggenerative models

Existing data pruning methods suffer significant performance degradation under high label noise and struggle to effectively retain informative samples. This work systematically investigates the behavior of pruning strategies in both noisy and noise-free settings, and for the first time explicitly identifies data redundancy, problematic samples, and inter-sample dependencies as three universal factors governing pruning efficacy. Through empirical analysis of two dominant pruning paradigms across standard classification benchmarks and mainstream neural architectures, the study demonstrates the consistent influence of these factors under diverse data distributions and training protocols. The findings not only expose fundamental limitations of current approaches but also offer a new perspective toward designing robust pruning methods.

data pruninglabel noiseproblematic samples

This study investigates the evolutionary dynamics of generative models trained iteratively on synthetic data contaminated with real data, aiming to mitigate model collapse induced by data pollution. Through statistical modeling, mixture distribution analysis, and theoretical analysis of iterative training dynamics—complemented by theoretical derivations and simulations based on next-token prediction language models—the work demonstrates that model collapse can be effectively avoided and the true data distribution even recovered, provided the mixture weight of real data remains non-zero over time and is paired with sufficient sample sizes. This mechanism consistently enhances performance across diverse model classes, offering both theoretical guarantees and practical guidance for sustainable iterative training.

contaminated sourcesgenerative modelsiterative training

Hot Scholars

QC

Qinyu Chen

Assistant Professor, Leiden University
Edge AIIC designNeuromorphic ComputingEvent-based vision
WY

Wenhan Yang

P.hD. student of Computer Science, University of California, Los Angeles
Self-supervised LearningModel Robustness
ZQ

Zhong-Qiu Wang

Associate Professor, Southern University of Science and Technology
Computer AuditionSpeech SeparationMicrophone ArrayAudio Signal Processing
KP

Kishan Panaganti

Tencent AI Lab
Large Reasoning ModelsReinforcement LearningRobust OptimizationStatistical Learning
ML

Mingjia Li

Beijing Institute of Technology
Generative ModelingDiffusion ModelsSemantic SegmentationDomain Adaptation/Generalization