Score
Designs and implements pipelines, frameworks, and tools to detect, mine, classify, and characterize failures and error cases—including failure-case mining (synthetic or real), failure detectors from internal signals, failure-mode classification, and root-cause labeling. Performs failure-mode analysis, fault localization, and error-pattern analysis to produce error-handling strategies, analysis workflows, and procedures for evaluating and transferring detectors across datasets.
A significant gap exists between academic research and industrial practice in debugging machine learning (ML) systems. Method: We propose the first comprehensive, lifecycle-spanning taxonomy of ML debugging faults and corresponding mitigation methods, derived from a systematic literature review (SLR), in-depth interviews with 28 ML practitioners, and empirical analysis of 1,247 GitHub issues. Contribution/Results: Our study identifies 13 core debugging challenges; only 48% are addressed by existing academic work, while 52.6% of GitHub issues and 70.3% of interview-elicited problems lack corresponding methodological support. Critically, we quantitatively demonstrate that over half of real-world ML debugging difficulties remain unaddressed by current research—revealing a substantial knowledge gap. This work establishes a foundational classification framework, provides empirical evidence of methodological coverage gaps, and delivers a prioritized roadmap to bridge the theory-practice divide in ML debugging.
Intermittent failures in continuous integration (CI) pipelines are notoriously difficult to diagnose, leading to wasted resources and reduced development efficiency. This work proposes FlaXifyer, a few-shot learning approach that integrates the interpretable AI technique LogSift to fine-tune pretrained language models on pipeline logs using only 12 labeled examples per failure class. The method simultaneously predicts failure categories and pinpoints critical log entries indicative of root causes. Evaluated on 2,458 real-world CI failures, FlaXifyer achieves a Macro F1 score of 84.3% and a Top-2 accuracy of 92.0%, reducing the required log inspection effort by 74.4%. Furthermore, it successfully identifies the underlying fault in 87% of cases, demonstrating its effectiveness in accelerating failure diagnosis with minimal labeled data.
This work addresses the challenges of multi-fault localization and GUI-component-level fault identification in software debugging. We propose a symbolic data mining approach that integrates Formal Concept Analysis (FCA) with association rule mining. By modeling test case outcomes (PASS/FAIL), GUI event sequences, and their mappings to underlying processors, our method automatically discovers fault patterns from coverage data. Crucially, we are the first to jointly extend both FCA and association rule mining to handle multi-fault scenarios and structured GUI event sequences—moving beyond conventional frequency-based fault localization paradigms. Experimental evaluation on benchmark programs including Trityp demonstrates that our approach significantly improves multi-fault detection rates and achieves higher precision in localizing faults at the GUI component level. The method is both empirically effective and scalable, supporting its applicability to complex, interactive software systems.
This work addresses the challenge of unreliable sensor readings in industrial inspection robots caused by occlusions, limited viewpoints, or environmental anomalies, which hinder real-time task status assessment. The authors propose a hybrid framework that integrates supervised fault classification with unsupervised anomaly detection, uniquely combining conformal prediction and world models to enable policy-agnostic, distribution-free early discrimination among three states—success, known faults, and out-of-distribution anomalies—using compressed video inputs. The approach facilitates training data quality evaluation and model feedback, achieving over 90% recognition accuracy on both office and industrial instrument inspection datasets. It outperforms human observers in decision speed and has been successfully deployed on a Boston Dynamics Spot robot for real-time operation.
Intermittent job failures in CI/CD—caused by non-deterministic factors such as environmental fluctuations or test flakiness—are frequently misclassified as code defects; existing retry-based heuristics suffer from high false-positive rates. This paper introduces the first few-shot learning framework for intermittent failure detection: leveraging only a small number of human-annotated logs (e.g., 12), it fine-tunes a compact language model to generate high-fidelity semantic embeddings and trains a lightweight classifier to distinguish failure types. By prioritizing annotation quality over quantity, the approach eliminates reliance on large-scale labeled datasets and mitigates misclassification stemming from heuristic policy discrepancies. Evaluated across multiple real-world projects, it achieves F1 scores of 70–88%, substantially outperforming state-of-the-art baselines (34–52%). Results demonstrate strong effectiveness, generalizability, and practical deployability in industrial CI/CD settings.
This work addresses the long-standing reliance on manual intervention in CI/CD pipeline failure management, which suffers from low efficiency and limited automation capabilities, particularly when handling unstructured failure information. To overcome these challenges, the authors propose an automated approach that integrates large language models (LLMs) with domain-specific knowledge—including historical failure data, pipeline context, and repair instructions—enabling end-to-end failure localization and repair generation for the first time in the large-scale industrial project SAP HANA. Experimental results demonstrate that, when augmented with domain knowledge, the method achieves a 97.4% accuracy in failure localization and generates fully correct repair solutions in 92.1% of cases, substantially outperforming knowledge-agnostic baselines. Ablation studies further confirm the critical contribution of historical failure data to overall performance gains.
This study addresses the limitations of traditional statistical fault localization (SFL), which relies solely on code execution traces and often fails to accurately pinpoint root causes. To overcome this, the authors systematically incorporate execution features—such as data flow, variable values, and branch conditions—extracted via the EFDD tool from the Tests4Py dataset. They train project-specific random forest models and map feature importance back to source code lines, integrating these insights with classical SFL formulas to enhance localization accuracy. Rigorous evaluation is conducted using a confounder-adjusted mixed-effects model and paired statistical tests. Experimental results demonstrate that the proposed approach significantly improves the accuracy of reference patches while reducing inspection effort at both line and function levels, confirming its robustness and practicality across multiple dimensions.
This work addresses the critical limitation in AI-driven automotive software fault analysis—the scarcity of representative fault datasets—by proposing a novel framework that integrates hardware-in-the-loop (HIL) simulation with real-time fault injection. For the first time, this approach enables synchronized generation of multimodal fault data, encompassing both time-series signals and textual logs, under both single-point and concurrent fault conditions. The framework effectively captures complex fault scenarios, yielding a comprehensive dataset that supports the training and validation of machine learning models. Empirical validation in a real-world testing environment demonstrates the framework’s practical applicability and effectiveness, offering a robust foundation for advancing fault diagnosis and resilience in automotive software systems.
This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).
This study addresses the lack of rigorous statistical assessment for the reliability of output structures in complex clustering pipelines that involve multiple data-dependent stages such as anomaly detection, feature selection, and clustering. To bridge this gap, the work systematically applies selective inference to the entire clustering analysis workflow, establishing a statistical framework that enables valid significance testing of final cluster assignments. The proposed method rigorously controls the type I error rate at any pre-specified nominal level and demonstrates strong empirical performance on both synthetic and real-world datasets. By doing so, it provides a principled and reliable foundation for statistical inference in multi-stage, data-driven clustering procedures.