detect and mitigate overfitting

Designs, builds, and analyzes workflows, evaluation protocols, and diagnostic tools that detect when models memorize training data or fit noise and that quantify the gap between training and held-out performance. Implements and tests mitigation strategies—including cross‑validation and other evaluation schemes, train/validation gap monitoring, leakage checks, regularization, data augmentation, controlled exposure to repeated examples, and pipeline changes—to prevent or reduce overfitting.

detectandmitigateoverfitting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.57
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the underexplored engineering challenges in existing machine learning evaluation frameworks, where operational issues and their root causes have lacked systematic investigation. To bridge this gap, the work formally establishes evaluation engineering as a distinct research direction within software engineering. Through an empirical analysis of 57 frameworks and a comprehensive categorization of 16,560 reported issues across a newly proposed five-stage workflow model, the study reveals that 41.4% of problems originate in the specification phase, while 61.7% of classified issues stem from missing functionality, inadequate documentation, and insufficient input validation. The findings yield a structured taxonomy of evaluation-related problems and provide empirical evidence to inform the design and improvement of robust evaluation systems.

empirical studyevaluation harnessesmachine learning

This work identifies an evaluation bias introduced by data augmentation (e.g., SMOTE, mutation-based augmentation) in scarce-data scenarios—particularly flaky test classification—where augmented samples inadvertently contaminate the test set, severely compromising fairness and reliability assessments. To address this, the authors first empirically identify and validate the critical phenomenon that “augmented data participation in testing” induces systematic evaluation distortion. They then propose a detection framework capable of disentangling training-induced bias from evaluation-induced bias, and design a bias-calibrated evaluation protocol. Experiments across multiple flaky-test benchmark datasets demonstrate that test sets containing augmented samples inflate accuracy by up to 23.7% and introduce F1-score deviations exceeding 0.15. This study establishes both theoretical foundations and practical guidelines for trustworthy model evaluation under data augmentation.

Bias in training and testing with augmented dataEvaluating augmented data effects in model testingImpact of data augmentation on model bias

MechDetect: Detecting Data-Dependent Errors

Dec 03, 2025
PJ
Philipp Jung
🏛️ Berlin University of Applied Sciences and Technology

The core challenge in data quality monitoring lies in error provenance—specifically, identifying the underlying mechanisms that generate errors—a problem largely overlooked by existing work, which seldom models such mechanisms explicitly. This paper focuses on errors arising from intrinsic dependencies within data and proposes MechDetect, the first method to systematically extend missing-data mechanism detection to diverse error types—including outliers, inconsistencies, and format violations. Leveraging joint statistical modeling and supervised learning, MechDetect simultaneously models tabular data and their error masks to automatically determine whether observed errors stem from inherent characteristics of the original data. Extensive experiments across multiple benchmark datasets demonstrate that MechDetect significantly outperforms state-of-the-art baselines in accurately diagnosing error-generation mechanisms. By providing mechanistic interpretability, it establishes a theoretical foundation and practical framework for explainable data repair.

Detect data-dependent error generation mechanismsEstimate error dependency using machine learning modelsExtend missing value analysis to other error types

Existing deep learning–based fault diagnosis methods suffer significant performance degradation on unseen programs, primarily due to a mismatch between conventional evaluation strategies and real-world deployment scenarios. This work introduces DynFault, a novel dataset comprising 38 real-world deep learning programs and 5,542 fault-injection traces, which for the first time systematically reveals and quantifies the performance gap—measured as a 0.190 drop in accuracy—between intra-program cross-validation and leave-one-program-out evaluation. The study identifies program-level feature structure as a key generalization bottleneck: curvature-based features prove effective for instability detection on unseen programs, whereas optimizer- and activation-related features exhibit predictive power only within seen programs, highlighting the heterogeneous generalizability of runtime features across programs.

deep learning programsevaluation gapfault diagnosis

A General Framework for Data-Use Auditing of ML Models

Jul 21, 2024
ZH
Zonghao Huang
🏛️ Duke University

To address copyright infringement and transparency concerns arising from unauthorized use of third-party data in machine learning model training, this paper proposes the first general-purpose, task-agnostic data usage auditing framework for black-box models. Methodologically, it innovatively integrates arbitrary black-box membership inference techniques with a custom sequential probability ratio test (SPRT), enabling zero assumptions about downstream tasks, strict control over false positive rates (tunable within 0.5%–5%), and cross-model generalization. The framework features a model-agnostic interface, supporting heterogeneous architectures including image classifiers and multimodal large language models. Extensive experiments on ImageNet classifiers and multimodal foundation models demonstrate an average detection accuracy exceeding 92%, with false positive rates consistently meeting user-specified thresholds. This work significantly enhances the quantifiability and reliability of training data provenance auditing.

Copyright IssuesData Usage TransparencyMachine Learning Model

Latest Papers

What's happening recently
View more

This study addresses the stagnation of enterprise AI initiatives in regulated financial institutions due to the absence of quantifiable evaluation criteria. Focusing on six document-intensive workflows, it systematically compares AI system performance across four model families and three tool configurations, distinguishing between demonstration and production environments. For the first time, it links deployment feasibility with human review rates. The authors propose a production-grade evaluation framework encompassing accuracy, reproducibility, traceability, and informative confidence, integrating multi-model comparison, confidence signals, source citation, and self-verification mechanisms. Experiments reveal that 56.1% of the 72 evaluated configurations meet production readiness thresholds. Incorporating source citation and confidence estimation reduces human review requirements to 49%, and adding self-verification further lowers this to 44%, albeit at the cost of reduced error tolerance.

AI deploymentconfidence calibrationproduction readiness

This work addresses a critical limitation in existing cloud skill testing, which focuses solely on task success rates and fails to reveal uncovered behaviors, leaving test adequacy unquantifiable. To bridge this gap, the study introduces the first formal definition of test coverage units, coverage relationships, and a computational pipeline for cloud skills. It proposes a coverage evaluation method grounded in natural language skill packages: by parsing user prompts and initial resource states, the approach reconstructs operational obligations and models workflow context to establish an end-to-end measurement pipeline. Integrating model-assisted candidate generation with expert review, the framework produces auditable coverage reports and closes the loop by mapping coverage gaps to source-level test improvement recommendations. Empirical results demonstrate that this methodology enables quantitative assessment of test coverage for production-grade cloud skills, substantially enhancing their reliability and observability.

AI AgentsCloud SkillsSkill Evaluation

Hot Scholars

LM

Leon Moonen

Full Professor, Simula Research Laboratory & BI Norwegian Business School
software engineeringsoftware securitysoftware analyticsdata mining & machine learning
MH

Max Hort

Simula Research Laboratory
Software EngineeringSoftware Fairness
ZS

Zhiwei Steven Wu

Carnegie Mellon University
Machine LearningDifferential PrivacyAlgorithmic FairnessGame Theory
PL

Percy Liang

Associate Professor of Computer Science, Stanford University
machine learningnatural language processing