Score
Designs and implements algorithms, sampling and weighting strategies, and training pipelines that identify, prioritize, or generate difficult, misclassified, rare, or negatively informative training examples (hard examples, hard negatives, bad cases, long‑tail samples) to improve model learning and robustness. This work includes mining heuristics, batch construction procedures, loss reweighting or curriculum schedules, negative‑pair selection for metric/contrastive losses, and diagnostic tools to analyze the impact of mined samples on performance and failure modes.
This work addresses the challenge of accurately predicting pretraining loss for large language models across varying model scales, batch sizes, and training steps—particularly under dynamically changing batch sizes and extreme extrapolation of compute budgets. To this end, the authors propose a loss prediction model grounded in a noisy quadratic system, which explicitly models test loss as a function of model size $N$, batch size $B$, and number of weight updates $K$. The framework enables joint optimization of training configurations under composite constraints on time, memory, and computational resources. Notably, it achieves high-precision loss prediction in variable-batch settings—a first in the field—and significantly outperforms existing heuristics such as Chinchilla, even when extrapolating up to 1000× beyond observed compute budgets. The recommended $(N, B, K)$ configurations closely align with empirically optimal solutions.
Machine learning models frequently suffer unexpected failures in real-world deployment, hindering practical adoption. Method: This paper introduces, for the first time, an orthogonal dichotomy framework distinguishing reliability from robustness, formally characterizing model failure mechanisms from first principles and systematically mapping them to engineering practices and real-world deployment scenarios. Our approach integrates probabilistic modeling, uncertainty quantification, adversarial robustness analysis, distributional shift detection, and system-level fault tree analysis—bridging theoretical insights with industrial-grade diagnostic tools and canonical failure case studies. Contribution/Results: We deliver an actionable failure attribution guide comprising rigorous theoretical foundations, an open-source toolchain, and cross-domain application exemplars. The framework significantly enhances model trustworthiness, debuggability, and deployment success rates.
The few-shot and imbalanced (S&I) learning problem suffers from severe generalization degradation and low interpretability due to scarce samples, extreme class imbalance, and ambiguous inter-class feature distributions. This paper proposes the first systematic analytical framework tailored to S&I learning, advocating that quantitative characterization of data properties—such as imbalance ratio and geometric complexity—must precede algorithmic design. The framework unifies multi-dimensional imbalance metrics, data complexity analysis, resampling strategies, classifier adaptation mechanisms, and an interpretable evaluation benchmark. Empirical evaluation on binary and multi-class extreme imbalance benchmarks reveals that classifier selection exerts significantly greater impact on performance than resampling improvements—exposing a fundamental flaw in prevailing heuristic-driven approaches. Our work establishes a theory-guided analytical paradigm and practical design principles for S&I learning, advancing both methodological rigor and empirical reproducibility.
Real-world structured data often suffer from demographic missingness, biased labels, and systematic sampling bias—yet existing robustness evaluations rely on random or simplistic corruptions, failing to expose worst-case vulnerabilities of high-risk ML systems. Method: We propose SAVAGE, the first causality-driven, black-box interpretable stress-testing framework for structured data. It models data dependencies via causal graphs and implements corruption templates to enable causal representation of structured data contamination. Its novel bilevel optimization algorithm supports end-to-end, targeted vulnerability discovery—even for pipelines containing non-differentiable components. Results: Experiments show that just 5% contamination generated by SAVAGE induces catastrophic performance drops, significantly outperforming baselines. Moreover, SAVAGE reveals that core assumptions underlying mainstream data cleaning and fairness-aware learning methods systematically fail under realistic data defects.
This work addresses the longstanding conflation in machine unlearning research between “untraining” and “unlearning,” which has led to ambiguous problem formulations and inadequate evaluation criteria. We formally distinguish these concepts for the first time: untraining aims to remove the influence of specific training samples, whereas true unlearning requires erasing the model’s knowledge of the entire underlying data distribution or concept those samples represent. Through theoretical formalization and a systematic review of existing literature, we establish a clear conceptual framework, reclassify current methods accordingly, and uncover critical challenges that have been overlooked. By clarifying foundational definitions, this study lays the groundwork for rigorous algorithmic evaluation, promotes standardization in the field, and delineates promising directions for future research.
This study addresses the limitations of existing approaches for automatically extracting machine learning (ML) pipeline structures, which often rely on manual annotations or suffer from insufficient generalization to keep pace with the rapid evolution of the ML ecosystem. The work presents the first systematic evaluation of small language models (SLMs) for reverse-engineering ML pipelines and proposes an SLM-based method for their automatic identification and reconstruction. Through comprehensive comparative experiments across multiple SLMs and rigorous statistical validation using Cochran’s Q, McNemar, and Pearson’s chi-squared tests, the authors demonstrate that the best-performing SLM significantly outperforms current methods and exhibits robustness across diverse classification schemes. This approach uncovers finer-grained patterns in data science practices and overcomes longstanding bottlenecks in scalability and domain adaptability inherent in traditional techniques.
Addressing model generalization under high annotation costs and data scarcity, this work establishes a unified theoretical and methodological framework for low-resource learning. Theoretically, it introduces the first agnostic active sampling theory integrated into the PAC learning framework, rigorously characterizing the trade-off between generalization error and labeling complexity. Methodologically, it proposes four novel optimization mechanisms—gradient-aware sampling, meta-iterative optimization, manifold-geometric modeling, and large-model-driven augmentation—and synergistically integrates transfer learning, reinforcement-based feedback, and hierarchical structural modeling. Empirical evaluation demonstrates that the proposed approach significantly enhances model robustness and generalization performance under limited annotations. This work provides both an interpretable, scalable theoretical foundation and a practical paradigm for data-constrained AI systems.
Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.
Transformer models exhibit fragile generalization on long sequences and structurally complex inputs. Method: We propose an automated curriculum learning framework driven by failure cases, centered on an executable-validator-guided counterexample discovery mechanism that dynamically generates challenging instances and constructs adaptive training curricula—without manual difficulty annotation. Our approach integrates counterexample-informed data augmentation, logical verification constraints, and Transformer fine-tuning to enable continuous self-correction. Contribution/Results: Experiments demonstrate a 30× improvement in sequence-length extrapolation on algorithmic reasoning and natural language tasks. Compared to uniform data augmentation, our method achieves a 3.75× speedup in computational efficiency and significantly outperforms both static training and conventional curriculum learning baselines.