Score
Design and build structured taxonomies and classification schemes that organize observable failure signatures and incidents into hierarchical and cross‑cutting classes, define labels and structured incident attributes, and specify recovery and fault‑model categories. Link those classes to suspected causes, enable failure‑rate measurement and probabilistic or qualitative analysis, and align taxonomy terms and priorities with regulatory or stakeholder requirements.
This work addresses the precision-recall trade-off in network intrusion detection, arising from the inherent diversity of cyber threats. Methodologically, we propose a novel detection framework that deeply integrates a cybersecurity incident taxonomy into the architectural design of detection networks. Guided by the taxonomy’s hierarchical semantic structure, we jointly leverage ontology-based analysis and controlled simulations to systematically identify the optimal operational equilibrium for detection strategies. Our key contribution is the first principled elevation of taxonomies from static labeling tools to structural priors embedded within detection models—explicitly encoding semantic relationships and evolutionary pathways among threat behaviors. Empirical evaluation across multiple public benchmark datasets demonstrates substantial improvements in holistic detection performance (average F1-score gain of 12.7%). Moreover, our analysis uncovers fundamental theoretical limits on detection set construction and establishes an interpretable pathway for performance optimization.
This work addresses the challenge of effectively reusing failure feedback from existing agent execution trajectories, which are often lengthy, instance-specific, and lack standardized failure descriptions. To overcome this, the authors propose an unsupervised method that automatically distills raw trajectories into a structured, evidence-backed failure taxonomy. This taxonomy forms an adaptive failure glossary organized along three axes—system-level, role-level, and domain-level—and serves as a unified feedback interface integrated into trajectory selection, runtime monitoring, and system search processes. Requiring no manual annotation, the glossary achieves a 10× compression ratio while exhibiting semantics closely aligned with expert annotations. Empirical results demonstrate significant performance gains across multiple benchmarks: SWE-agent’s resolution rate improves from 60% to 70%, Claude Code reaches 70.7%, and Terminal-Bench 2.0 accuracy increases by 8–15 points.
In complex equipment fault diagnosis, domain expertise is difficult to formalize, and manual fault tree construction is inefficient. Method: This paper proposes a knowledge graph–based approach for automated fault tree synthesis. It introduces a lightweight, semantically rich knowledge graph representation that enables semi-automatic extraction of failure logic relationships from unstructured documents (e.g., maintenance manuals) and structured/functional models. Leveraging hierarchical modeling and semantic reasoning, the method generates fault trees in a fully structured manner—without requiring historical fault data, relying solely on engineering knowledge. Contribution/Results: The synthesized fault trees are inherently interpretable and accurately capture system-level failure propagation paths. Experimental validation on the Lycoming O-320 aircraft engine demonstrates substantial improvements in diagnostic modeling efficiency and engineering applicability.
Addressing the challenge of multi-label automatic annotation for large-scale, hierarchical classification systems in software requirements engineering, this study proposes a sentence-level zero-shot classification paradigm to circumvent the high annotation costs associated with supervised training. We introduce the first industrial-scale requirements annotation benchmark comprising 769 taxonomy labels and systematically demonstrate a strong negative correlation between the number of taxonomy leaf nodes and classification recall. We further propose a zero-shot multi-label classification method leveraging SBERT sentence embeddings, achieving significant improvements in recall. Empirical evaluation reveals that hierarchical strategies yield no consistent performance gain across settings. Our work validates the effectiveness and feasibility of zero-shot learning for large-scale requirements classification, offering a scalable, low-human-effort automation solution for requirements tracing. (138 words)
To address the challenge of manual diagnosis for anomalous trace faults in microservice systems, this paper proposes TraFaultDia—a novel framework that establishes, for the first time, a multi-system, cross-domain few-shot anomaly trace classification paradigm. TraFaultDia integrates Model-Agnostic Meta-Learning (MAML), Graph Neural Network (GNN)-based trace representation learning, multi-task few-shot classification, and a cross-system fault pattern alignment mechanism. It achieves high-accuracy fault type identification on unseen systems using only ten labeled samples per class. Extensive experiments on TrainTicket and OnlineBoutique demonstrate that TraFaultDia attains an average accuracy of 93.26% (92.19% on novel tasks) in intra-system settings and 85.20% (84.77% on novel tasks) in cross-system transfer scenarios—significantly enabling zero-effort localization of faulty components and root causes.
Current agent evaluation practices often reduce failures to system-level outcomes, making it difficult to pinpoint root causes or guide effective remediation. This work proposes an interaction-centric failure taxonomy and introduces, for the first time, a cross-architectural and generalizable framework for failure localization. The framework maps 41 distinct failure modes onto interaction edges between components—such as models, toolchains, and environments—and explicitly delineates responsibility boundaries among them. By integrating component interaction graph attribution, multi-source trajectory analysis, and an independent reasoning agent-based evaluator, the approach enables reproducible validation. Experiments across four state-of-the-art models demonstrate that the strongest evaluator achieves a Cohen’s κ of 0.76 with human annotations, confirming the taxonomy’s generalizability and consensus alignment.
This study addresses the lack of a systematic taxonomy for failure modes in multi-provider large language model (LLM) service gateways, which hinders effective detection and diagnosis in production environments. The work proposes the first dual-axis structured failure classification framework for LLM gateways, categorizing failures by origin layer—spanning network/transport, streaming/protocol, state/session, model behavior, and governance/cost—and by detectability (explicit vs. implicit). Through root cause analysis, stress testing, mining of public bug reports, and protocol-level debugging, the authors construct a catalog of five validated failure cases, three of which include reproducible scripts. Notably, the study uncovers two previously undocumented silent failures: conversation history loss due to concurrency races and stream-index collisions corrupting tool-call payloads. These return HTTP 200 responses and pass standard health checks, evading detection without semantic-level observability, thereby posing significant threats to application reliability.
This study addresses the challenge of analyzing unstructured maintenance logs from wind turbines, which hinder quantitative reliability assessment. The authors propose a model-agnostic large language model (LLM) framework that enables fully automated semantic parsing and structuring of wind turbine logs for the first time. By integrating semantic extraction, domain-specific lexicon construction, and encoding correction, the method automatically derives an evidence-based taxonomy of repair actions and failure modes. Applied to 16,316 log entries, the approach successfully structures over 70% of the data, substantially correcting misclassifications, recovering missing codes, and reducing the subjectivity inherent in traditional Failure Mode and Effects Analysis (FMEA). Furthermore, it establishes a scalable pipeline for cross-turbine reliability knowledge integration and metric generation.
This work addresses the limitations of generic pretrained models in industrial-scale video and live-stream content moderation, where platform-specific data distributions, policy objectives, and safety constraints are poorly aligned with off-the-shelf solutions, and systematic failure diagnosis and remediation mechanisms are lacking. The paper introduces a diagnostic methodology for audio-visual language models (AVLMs) that pioneers the characterization of model failures through observable feature signatures and establishes a principled mapping between failure categories and targeted intervention strategies, replacing heuristic trial-and-error approaches. Built upon multimodal foundation model architectures and validated on real-world platform traffic, this framework enables precise interventions throughout the model lifecycle. The resulting AVLM system has been deployed across more than 100 regions globally, significantly improving moderation accuracy and traceability for high-noise, semantically ambiguous, and highly diverse content.