error analysis

Diagnosing failure modes and performance variation through qualitative and quantitative examination of errors across domains, inputs, or model behaviors; used to identify robustness issues, sources of silent misclassification, and domain-specific weaknesses.

erroranalysis

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Towards transparent and data-driven fault detection in manufacturing: A case study on univariate, discrete time series

Jun 30, 2025
BH
Bernd Hofmann
🏛️ Friedrich-Alexander-Universität Erlangen-Nürnberg

To address the trade-off between poor adaptability of conventional methods and weak interpretability of data-driven models in manufacturing fault detection, this paper proposes an interpretable fault detection framework tailored to univariate discrete-time series from crimping processes. Methodologically, it integrates supervised multi-class classification, post-hoc Shapley value explanation, and domain-specific visualization to jointly deliver fault classification and human-readable decision rationale. Its key contribution lies in a human-centered explanation mapping mechanism, rigorously validated through quantitative perturbation analysis and expert evaluation. Experimental results demonstrate a classification accuracy of 95.9%, with explanations exhibiting both statistical relevance and operational readability. The framework significantly enhances the reliability, trustworthiness, and practical utility of industrial quality control systems.

Enhancing interpretability of machine learning for industrial quality controlEnsuring product quality in safety-critical manufacturing applicationsOvercoming black-box limitations in data-driven fault detection models

From PREVENTion to REACTion: Enhancing Failure Resolution in Naval Systems

Aug 21, 2025
MT
Maria Teresa Rossi
🏛️ University of Milano -Bicocca

Naval systems frequently exhibit anomalous behaviors due to wear, misuse, or component failures—challenges that hinder timely detection and precise remediation. To address this, we propose a predictive-diagnostic closed-loop framework that tightly integrates the existing failure prediction system PREVENT with a newly designed responsive troubleshooting module, REACT. Methodologically, the framework synergizes multi-source time-series anomaly detection with domain-knowledge-driven fault-isolation process modeling, enabling end-to-end automation—from anomaly alerting and root-cause localization to actionable remediation recommendations. Evaluated on operational shipboard systems deployed by Fincantieri, the framework reduces mean time to fault localization by 42%, significantly improves operational response efficiency, and demonstrates strong generalizability across diverse industrial domains.

Enhancing failure detection and resolution in naval systemsExtending predictive maintenance to industrial productsIntegrating anomaly detection with troubleshooting procedures

To address poor generalization in fault detection under dynamic industrial operating conditions—caused by train-test distribution shift—this paper proposes TARD, a novel test-time domain adaptation method operating continuously during inference. TARD introduces a pioneering dual-branch architecture that decouples system parameters from sensor measurements, applying distinct adaptive strategies to each: dynamic batch normalization for the parameter branch, and unsupervised domain alignment coupled with temporal feature disentanglement for the sensor branch. The method enables lightweight, online model updating without requiring labeled target-domain data. Evaluated on two real-world multiphase flow systems, TARD significantly improves early fault detection accuracy and robustness compared to state-of-the-art domain adaptation approaches. It effectively mitigates distribution shifts arising from both data scarcity and evolving operational conditions, demonstrating strong practical applicability in industrial monitoring scenarios.

Address distribution shifts between training and testing dataDetect faults under evolving industrial operating conditionsImprove fault detection with limited training data representativeness

Quantitative Measurement of Cyber Resilience: Modeling and Experimentation

Mar 28, 2023
MJ
Michael J. Weisman
🏛️ DEVCOM Army Research Laboratory | Pennsylvania State University | ICF International | University of California, Irvine

Current cyber-physical systems (CPS) in vehicular environments lack quantitative, experimentally grounded methods for assessing network resilience. Method: This study constructs an experimental testbed replicating real-world truck operational conditions and conducts multiple rounds of malware injection attacks, simultaneously collecting network- and physical-layer data on resistance and recovery behaviors. Contribution/Results: We introduce the novel concept of “bonware” to holistically characterize both cybersecurity defense capability and physical resilience, formalized via an analytically tractable mathematical model. We further define and extract experimentally identifiable, quantitative resilience metrics—termed elastic features—for the first time. Sensitivity analysis confirms these metrics exhibit significant discriminability with respect to attack intensity, defensive strategies, and physical redundancy. This work bridges a critical gap by advancing vehicular CPS resilience from qualitative description to quantifiable, comparable, and optimizable measurement.

Attack RecoveryCyber ResilienceMeasurement Tools

A General Approach for Determining Applicability Domain of Machine Learning Models

May 28, 2024
LE
Lane E. Schultz
🏛️ University of Wisconsin-Madison | Carnegie Mellon University

This study addresses the lack of generality and interpretability in applicability domain (AD) estimation for machine learning models. We propose a unified AD assessment framework based on kernel density estimation (KDE), which quantifies the distance of a query sample from the training data distribution in feature space and establishes a quantitative relationship among distance, prediction error, and uncertainty. Chemical prior knowledge is incorporated to calibrate the AD decision threshold. To our knowledge, this is the first method enabling consistent, cross-model and cross-task AD evaluation across diverse models—including random forests (RF), gradient-boosted decision trees (GBDT), and graph neural networks (GNN)—and heterogeneous materials datasets (crystals, molecules, alloys). Experiments demonstrate that large KDE-derived distances strongly correlate with high prediction residuals and elevated uncertainty estimates. An open-source toolkit enables automated in-domain/out-of-domain classification. The implementation and documentation are publicly available.

Assessing data distance in feature space for domain determinationDetermining applicability domain of machine learning modelsIdentifying in-domain versus out-of-domain predictions reliably

Latest Papers

What's happening recently
View more

This study addresses the reliability of probabilistic uncertainty quantification (UQ) in software defect prediction, particularly its ability to reflect model performance and calibration—especially in cross-project settings, where systematic validation remains lacking. Through a large-scale empirical analysis of 16 classifiers across 36 within-project and 32 cross-project datasets, the work examines the relationships between five UQ metrics and six performance measures alongside three calibration metrics. It reveals, for the first time, a strong context dependency: within-project, UQ correlates strongly with false positive rate and AUC, but these correlations substantially weaken or even reverse in cross-project scenarios. Notably, high-performing models can still exhibit severe miscalibration. These findings indicate that UQ signals are not directly transferable and must be evaluated independently relative to specific objectives, using multidimensional calibration assessments.

CalibrationCross-Project PredictionPerformance Evaluation

This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).

Agentic SystemsMonitoringStructural Defects

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

This work addresses the challenge of detecting and diagnosing training anomalies in deep learning systems, which often stem from subtle implementation defects yet lack labeled training trajectory data for systematic study. To bridge this gap, the authors construct Deep4ge, a large-scale benchmark dataset derived from 59 real-world TensorFlow/Keras programs. By applying 27 source-level transformations, they inject seven representative fault types, yielding 14,227 training runs—comprising 9,845 faulty and 4,382 normal executions. Each run is annotated with 26 fine-grained features spanning weights, gradients, activations, loss, accuracy, learning rate, and hardware utilization, along with four evaluation metrics. Deep4ge is the first publicly available, controlled dataset of DNN training trajectories with explicit fault labels, enabling research in binary anomaly detection, multi-class fault diagnosis, and early prediction. The dataset and fault-injection framework are open-sourced.

datasetdeep learningfault detection

This study addresses the high fine-tuning costs and associated risks—such as knowledge degradation, diminished instruction-following capability, and increased hallucination—faced by small-scale large language models in cybersecurity question-answering tasks. To mitigate these issues, the authors propose the FiT diagnostic framework, which evaluates model suitability prior to fine-tuning along three dimensions: lexical recognition, parametric knowledge, and contextualization of retrieved information. The work establishes the first task-oriented diagnostic system tailored for cybersecurity QA, integrating knowledge-focused and instruction-focused fine-tuning paradigms with retrieval-augmented evaluation. Through empirical analysis of five open-source 7B models, the study reveals that fine-tuning generally impairs lexical and parametric knowledge: knowledge-focused fine-tuning induces mild yet consistently ranked degradation, whereas instruction-focused fine-tuning triggers severe knowledge collapse while preserving the ability to contextualize retrieved information. Notably, FiT scores effectively predict post-fine-tuning performance trends.

cybersecurity QAdiagnostic evaluationfine-tuning

Hot Scholars

ZZ

Zibin Zheng

IEEE Fellow, Highly Cited Researcher, Sun Yat-sen University, China
BlockchainSmart ContractServices ComputingSoftware Reliability
CT

Christoph Treude

Associate Professor of Computer Science, Singapore Management University
Software EngineeringEmpirical Software EngineeringHuman-AI InteractionAI for Science
DS

Dawn Song

Professor of Computer Science, UC Berkeley
Computer Security and Privacy
AE

Ahmed E. Hassan

Mustafa Prize Laureate, ACM/IEEE/NSERC Steacie Fellow, ACM Influential/IEEE Distinguished Educator
Mining Software RepositoriesSoftware AnalyticsEmpirical Software EngineeringSoftware
FK

Foutse Khomh

NSERC Arthur B. McDonald Fellow, CRC Tier 1, Canada CIFAR AI Chair, FRQ-IVADO Chair, Full Professor
Software engineeringMachine learning systems engineeringMining software repositoriesReverse