Score
Design and implement training procedures, curricula, and data pipelines that combine clean and noisy examples into a single training process, including sampling strategies, loss weighting, and augmentation so models are trained on mixed-quality data. Build experiments and analyses that measure how such combined clean–noisy training affects model robustness, regularization, and architecture-dependent performance.
To address inefficiencies in manual dataset management and challenges in optimizing mixed-data sampling strategies and ordering amid rapidly expanding multi-source datasets for large model training, this paper proposes the first declarative, centralized, read-only data plane architecture. The architecture decouples data management from training frameworks and enables metadata-driven, cross-attribute (e.g., language, source) dynamic data mixing. It natively integrates state-of-the-art algorithms such as Adaptive Data Optimization (ADO) and supports real-time adjustment of mixing policies via model feedback. The system employs a training-framework-agnostic abstraction layer, a distributed sampling engine, and hardware-aware optimizations for GH200 superchip clusters. Empirical evaluation on 256 GPUs demonstrates zero throughput bottlenecks, substantial improvements in LLM and VLM performance, and a tenfold reduction in experimental iteration cost for mixture strategy tuning.
Dataset mixing for large language model fine-tuning typically relies on labor-intensive trial-and-error and incurs substantial computational overhead. Method: This paper proposes a zero-shot dataset composition selection method that leverages model merging as a proxy evaluator—specifically, employing weighted averaging and task vector fusion to predict downstream performance of candidate dataset mixtures, while jointly optimizing mixture weights. Unlike conventional heuristic strategies requiring repeated full fine-tuning, our approach eliminates the need for any fine-tuning during evaluation. Contribution/Results: Experiments across multiple benchmarks demonstrate that our method significantly outperforms existing dataset selection techniques. It achieves comparable or improved final fine-tuned model performance while reducing computational cost by approximately 70% in GPU-hours, enabling efficient, scalable dataset composition design.
In noisy label learning, sample selection suffers from dual biases: data bias (imbalanced selection sets) and training bias (error accumulation). To address these issues, this paper proposes ITEM, a noise-tolerant expert model. ITEM is the first to jointly model and mitigate both biases in a unified framework. It introduces a lightweight multi-expert robust network architecture, integrated with a dual-weighted class-discriminative sampler and a hybrid mini-batch training strategy. Additionally, an error-robust optimization mechanism is incorporated to enhance generalization under label noise. Extensive experiments on multiple benchmark datasets with synthetic and real-world label noise demonstrate that ITEM consistently outperforms state-of-the-art methods, achieving average accuracy gains of 3–5% while reducing parameter count by over 20%. The source code is publicly available.
This work investigates how label noise in pretraining data affects the generalization of foundation models, revealing that while such noise may improve in-distribution (ID) performance, it inevitably degrades out-of-distribution (OOD) generalization—primarily by distorting the feature space geometry. To address this, the authors introduce “noisy-model tuning” as a novel paradigm and propose NMTune, a generic, parameter-efficient feature-space calibration method applicable to both white-box and black-box models. Extensive evaluation across synthetic and real-world noisy datasets—including ImageNet-1K, YFCC15M, and CC12M—covers diverse pretraining paradigms (fully supervised and vision-language contrastive), model architectures, downstream tasks, and tuning strategies. Results demonstrate that NMTune consistently mitigates noise-induced degradation, significantly improving OOD generalization across vision and language models—including proprietary API-based models—without dependence on model scale or task specificity.
To address overfitting and poor noise robustness in large-model fine-tuning under few-shot settings, this paper proposes a novel fine-tuning framework that jointly enhances generalization and robustness. Methodologically, it introduces the first layer-wise L2-distance constraint regularization, integrated with confidence-guided self-label correction and dynamic reweighting. Theoretically, it is the first to systematically incorporate PAC-Bayes generalization bounds and noise stability analysis into fine-tuning design. Extensive experiments across seven image classification benchmarks demonstrate that the framework achieves an average accuracy improvement of 1.76%, gains +0.75% under few-shot conditions, and significantly outperforms baselines by +3.56% under label noise—validating both its effectiveness and robustness.
This work systematically investigates the impact of loss weighting strategies and output parameterizations on model performance in flow matching. Through numerical experiments on both synthetic data with controllable geometric structures and real-world images, the study disentangles their interaction effects across varying data manifold dimensions, model architectures, and dataset scales, using PSNR and FID as evaluation metrics. The analysis reveals, for the first time, how the optimal choice of loss weighting and parameterization depends critically on the intrinsic structure of the data. Building on these insights, the authors formulate practical design principles that substantially improve denoising accuracy and generation quality.
Existing data pruning methods suffer significant performance degradation under high label noise and struggle to effectively retain informative samples. This work systematically investigates the behavior of pruning strategies in both noisy and noise-free settings, and for the first time explicitly identifies data redundancy, problematic samples, and inter-sample dependencies as three universal factors governing pruning efficacy. Through empirical analysis of two dominant pruning paradigms across standard classification benchmarks and mainstream neural architectures, the study demonstrates the consistent influence of these factors under diverse data distributions and training protocols. The findings not only expose fundamental limitations of current approaches but also offer a new perspective toward designing robust pruning methods.
研究通过系统性调整学习率、批次大小等参数,对比LoRA和全微调方法,在不同模型和数据集上优化监督微调效果。
本文探讨了在数据不完美条件下的机器学习挑战,并通过信息损失、经验风险偏差等机制组织代表性方法,如重建生成、再平衡与表示校准等。
This study investigates the evolutionary dynamics of generative models trained iteratively on synthetic data contaminated with real data, aiming to mitigate model collapse induced by data pollution. Through statistical modeling, mixture distribution analysis, and theoretical analysis of iterative training dynamics—complemented by theoretical derivations and simulations based on next-token prediction language models—the work demonstrates that model collapse can be effectively avoided and the true data distribution even recovered, provided the mixture weight of real data remains non-zero over time and is paired with sufficient sample sizes. This mechanism consistently enhances performance across diverse model classes, offering both theoretical guarantees and practical guidance for sustainable iterative training.