Score
Designs and evaluates techniques that augment spectrum-based fault localization by extracting and using execution-derived features to produce ranked lists of suspicious program elements. This competence includes building predictive models (e.g., per-subject random forests) on execution features, deriving feature importances, and combining those importance weights with SFL scores to re-rank and prioritize lines or components for debugging.
This study addresses the limitations of traditional statistical fault localization (SFL), which relies solely on code execution traces and often fails to accurately pinpoint root causes. To overcome this, the authors systematically incorporate execution features—such as data flow, variable values, and branch conditions—extracted via the EFDD tool from the Tests4Py dataset. They train project-specific random forest models and map feature importance back to source code lines, integrating these insights with classical SFL formulas to enhance localization accuracy. Rigorous evaluation is conducted using a confounder-adjusted mixed-effects model and paired statistical tests. Experimental results demonstrate that the proposed approach significantly improves the accuracy of reference patches while reducing inspection effort at both line and function levels, confirming its robustness and practicality across multiple dimensions.
Traditional spectrum-based fault localization (SBFL) relies solely on coverage spectra across test cases, ignoring rich runtime, control-flow, and lexical information embedded in program execution traces. This work proposes a lightweight context-aware SBFL method that requires no GPU or heavyweight program analysis. By augmenting execution traces with multi-granularity features—including variable values, branch outcomes, and abstract syntax tree (AST) node encodings—and integrating them via machine learning models, the approach enables context-sensitive suspiciousness re-ranking. Its core contribution lies in a low-overhead feature extraction and fusion mechanism that significantly improves localization accuracy without modifying the original spectrum input. Evaluated on QuixBugs and competitive programming benchmarks, the method consistently outperforms classical formulas (e.g., Ochiai), achieving an average 12.7% improvement in Top-1 localization accuracy. It thus offers high precision, minimal computational overhead, and strong practical applicability.
This paper addresses the dual challenges of low fault localization accuracy and weak root-cause interpretability in software debugging. To this end, we propose an interpretable diagnosis method based on multi-execution feature fusion. Through empirical analysis of 310 real-world defects, we first establish—systematically and for the first time—that scalar pairs constitute the strongest failure-correlated features. Building upon this insight, we design a joint modeling framework that integrates 17 fine-grained execution features, including variable values, branch conditions, and definition-use chains. We further develop a feature-importance-driven interpretable decision tree model that automatically generates human-readable diagnostic rules. Evaluation across 20 open-source projects demonstrates that our approach significantly improves both fault localization accuracy and root-cause identification depth, substantially reducing developer debugging time. The method achieves a favorable balance between high precision and strong interpretability.
Spectrum-Based Fault Localization (SBFL) fails when no failing tests are available to trigger faults. Method: This paper systematically demonstrates, for the first time, that stack traces from crash reports can serve as pseudo-failure signals in lieu of actual failing tests, and proposes SBEST—a novel SBFL method that integrates exception-location semantics with method-call-graph reachability to embed stack-trace information into the spectrum analysis framework. SBEST jointly leverages test coverage matrices and parsed stack traces to enable precise fault localization even in the absence of failing tests. Results: Experiments show SBEST improves Mean Average Precision (MAP) by 32.22% and Mean Reciprocal Rank (MRR) by 17.43% over the baseline MAP method. Moreover, 98.3% of defect-fixing intentions align with stack-trace anomalies, and 78.3% of defective methods are reachable within an average of 0.34 call-graph hops. This work establishes a new lightweight, crash-driven paradigm for fault localization.
Existing formula-based fault localization (FBFL) methods for multi-fault C programs struggle to simultaneously ensure cross-failure-test consistency and subset-minimality of diagnoses. To address this, we propose a MaxSAT-based approach that integrates model-based diagnosis (MBD) with multi-observation aggregation. Our method employs static analysis to extract program semantics and constructs a unified diagnostic model covering multiple failing test cases; MaxSAT solving then guarantees that the resulting diagnoses are both subset-minimal and consistent across all failing tests. This work is the first to systematically apply the MBD paradigm to multi-fault localization in C programs, eliminating redundant diagnoses. Evaluated on the TCAS and C-Pack-IPAs benchmarks, our approach achieves higher fault-localization accuracy than BugAssist and SNIPER, while demonstrating significantly improved runtime efficiency.
This work proposes a novel integration of Delta Debugging (DDMIN) with spectrum-based fault localization (SBFL) to precisely identify faulty statements using only a single failing input. While traditional DDMIN effectively minimizes failing inputs, it does not pinpoint the exact fault location. The proposed approach leverages the diverse passing and failing test cases automatically generated during the DDMIN reduction process to compute and rank statement suspiciousness via SBFL techniques such as Jaccard. Evaluated on 136 programs from QuixBugs and Codeflaws, the method consistently ranks the actual faulty statement within the top three in most cases and requires inspecting fewer than 20% of executable lines, significantly improving both the efficiency and accuracy of fault localization.
While developers commonly identify bug-introducing commits (BICs) via binary search over version histories to localize compiler bugs, existing spectrum-based fault localization (SBFL) techniques have not been systematically benchmarked against this simple yet widely adopted strategy. Method: This paper introduces the BIC-localization strategy (“Basic”) as a baseline and conducts a comparative evaluation against state-of-the-art SBFL methods on 60 real-world defects each in GCC and LLVM. BICs are localized via binary search guided by version history analysis and SBFL-based bug reproduction; Top-1 and Top-5 localization accuracy are quantified. Contribution/Results: Basic matches or surpasses advanced SBFL techniques across most scenarios, exposing critical limitations in current SBFL evaluation paradigms. The study establishes BIC localization as a new practical benchmark for compiler defect localization and provides methodological insights for evaluating the real-world utility of fault localization techniques.
Traditional fault localization approaches struggle to handle semantic errors, while existing large language model (LLM)-based methods often produce stochastic, unverifiable outputs that conflate root causes with cascading effects. This work proposes SemLoc, a novel framework that introduces structured semantic grounding for LLM-based reasoning: it anchors free-form LLM-generated explanations to program-specific reference points, constructs a semantic violation spectrum via dynamic instrumentation, and incorporates a counterfactual verification mechanism to identify critical causal constraints. The approach enables runtime validation and cross-test attribution, achieving a Top-1 accuracy of 42.8% (Top-3: 68%) on the SemFault-250 benchmark while inspecting only 7.6% of code lines; ablation studies show that counterfactual verification contributes a 12% absolute gain in accuracy.
Test Code Fault Localization (TCFL) in continuous integration is hindered by black-box environments, sparse diagnostic signals, and vast search spaces, limiting the effectiveness of existing approaches. This work proposes SPARK, a novel framework that synergistically integrates historical debugging knowledge with large language models (LLMs). SPARK employs a retrieval-augmented mechanism to identify historically similar failing test cases from CI logs and applies selective line-level annotations to guide the LLM toward suspicious code regions. This strategy avoids prompt length explosion while significantly improving localization accuracy, particularly in complex scenarios involving multiple faults. Experimental results demonstrate that SPARK outperforms current LLM-based TCFL methods on three industrial-scale Python test datasets, achieving higher precision in identifying multiple fault locations while maintaining manageable inference overhead and token consumption.
This work addresses the challenge of precise root cause identification and fine-grained localization in Linux kernel fault diagnosis, which is hindered by insufficient deep modeling of kernel-specific characteristics. To overcome this limitation, we propose CoHiKer, a novel approach leveraging large language models to perform contrastive reasoning by analyzing behavioral discrepancies between failing and passing test cases. CoHiKer integrates multi-level contextual information—including crash reports, system call semantics, and cross-file dependencies—into a hierarchical analysis framework that systematically narrows down the fault localization scope. Experimental results demonstrate that CoHiKer achieves significant improvements of 26.07% and 56.85% in Top-1 accuracy at the file and method levels, respectively, on kernel datasets, while substantially reducing token consumption. Furthermore, it exhibits strong generalization capability on non-kernel datasets.