Score
Designs and builds diagnostic datasets that collect and organize expert-labeled failure cases and fine-grained causal annotations, structuring data into modular components that isolate fault categories and causal factors. Defines sampling and labeling protocols to ensure coverage of diverse faults and systems (including AIOps/root-cause scenarios) and establishes validation and benchmarking procedures to verify dataset quality and suitability for model testing and analysis.
Large-scale, heterogeneous software defect datasets hinder efficient navigation and reuse by researchers. This paper systematically surveys 132 publicly available defect datasets and proposes a multidimensional evaluation framework—covering domain coverage, defect types, programming language distribution, construction methodologies, accessibility, and citation contexts—to achieve the first standardized metadata harmonization and empirical usability validation across a hundred-plus datasets. Through bibliometric analysis, citation network mapping, and cross-dimensional clustering, we identify test generation and automated program repair as the most widely supported application domains, while critical defect categories—including concurrency and security vulnerabilities—remain severely underrepresented. We further propose actionable dataset curation guidelines and a reusable assessment template, diagnosing pervasive issues such as incomplete coverage, disorganized structure, and outdated maintenance. Our work establishes a robust, empirically grounded data foundation for software defect detection, localization, repair, and AI-driven development.
In complex equipment fault diagnosis, domain expertise is difficult to formalize, and manual fault tree construction is inefficient. Method: This paper proposes a knowledge graph–based approach for automated fault tree synthesis. It introduces a lightweight, semantically rich knowledge graph representation that enables semi-automatic extraction of failure logic relationships from unstructured documents (e.g., maintenance manuals) and structured/functional models. Leveraging hierarchical modeling and semantic reasoning, the method generates fault trees in a fully structured manner—without requiring historical fault data, relying solely on engineering knowledge. Contribution/Results: The synthesized fault trees are inherently interpretable and accurately capture system-level failure propagation paths. Experimental validation on the Lycoming O-320 aircraft engine demonstrates substantial improvements in diagnostic modeling efficiency and engineering applicability.
The root cause analysis (RCA) community suffers from a critical shortage of large-scale, open-source, multimodal benchmark datasets, hindering rigorous method evaluation and advancement. To address this, we introduce LEMMA-RCA—the first cross-domain, multimodal, open-source RCA dataset tailored for IT/OT systems, encompassing realistic failure scenarios from microservices, water supply, and wastewater treatment. It comprises hundreds of system entities and fine-grained causal relationship annotations. Uniquely integrating four heterogeneous modalities—time-series metrics, logs, topology graphs, and alerts—it supports both offline/online and unimodal/multimodal RCA evaluation. Built via distributed monitoring, multi-source alignment, controllable fault injection, and causal graph annotation, the dataset ensures high fidelity and representativeness. Extensive evaluation across eight baseline methods demonstrates that multimodal joint modeling improves average F1-score by 23.6%. LEMMA-RCA is publicly released, establishing a new community benchmark for RCA research.
Existing benchmarks for microservice fault diagnosis focus solely on final answers, overlooking the systematic reasoning processes of large language model agents. This work proposes the first evaluation paradigm centered on the diagnostic reasoning process, introducing AIOps2025 and RCA100—large-scale, expert-annotated datasets encompassing three key dimensions: fault localization, identification, and root cause attribution. These datasets integrate multimodal observability data with causal reasoning analysis and cover over 500 real-world failure cases. The benchmark’s effectiveness has been validated through an international competition involving more than 6,000 teams, establishing it as the first reasoning-oriented benchmark for intelligent microservice fault diagnosis to be empirically validated at scale.
This work proposes an intelligent agent-based diagnostic framework leveraging large language models (LLMs) to overcome the limitations of traditional root cause analysis methods, which rely on hard-coded rules, incur high maintenance costs, and are tightly coupled with infrastructure. By integrating a Model Context Protocol (MCP) and a constrained tool space, the framework enables agents to autonomously invoke tools for service querying, dependency retrieval, and multi-source data analysis, facilitating stepwise reasoning to pinpoint root causes. A structured investigation protocol ensures traceable and reproducible inference while maintaining robustness under incomplete or ambiguous information, effectively decoupling the model from underlying infrastructure. This approach lays the foundation for autonomous fault diagnosis and change impact assessment, paving the way for automated remediation and risk prediction, thereby significantly enhancing operational efficiency and system safety.
This work addresses the critical limitation in AI-driven automotive software fault analysis—the scarcity of representative fault datasets—by proposing a novel framework that integrates hardware-in-the-loop (HIL) simulation with real-time fault injection. For the first time, this approach enables synchronized generation of multimodal fault data, encompassing both time-series signals and textual logs, under both single-point and concurrent fault conditions. The framework effectively captures complex fault scenarios, yielding a comprehensive dataset that supports the training and validation of machine learning models. Empirical validation in a real-world testing environment demonstrates the framework’s practical applicability and effectiveness, offering a robust foundation for advancing fault diagnosis and resilience in automotive software systems.
This work addresses a critical limitation in current AI4Science practices, which often treat datasets as static interfaces while neglecting the uncertainties and implicit assumptions introduced by the multi-stage processing pipeline from raw measurements to curated datasets. To remedy this, the paper proposes a “computable observation framework” that explicitly models this pipeline as an auditable and reproducible inference component, capturing its configuration, validity, and associated uncertainties. By integrating scientific workflow analysis, uncertainty quantification, and cross-dataset stability assessment, the framework enables the construction of domain-specific observation protocols. Empirical evaluation on large-scale neuroscience data reveals that only approximately 0.0004% of processing pipelines exhibit cross-dataset stability, exposing severe fragility in current practices and underscoring the framework’s essential role in uncovering hidden assumptions, validating transferability, and controlling for multiplicity.