Score
Designs and builds automated diagnostic systems, tooling, and scripts that detect, isolate, and report failures across software, hardware, embedded/on-device components, and databases. Implements runtime and remote instrumentation, cross-layer correlation and failure root‑cause analysis, and AI‑assisted or creatively heuristics-driven diagnostics, including hardware diagnostic tooling, SQL diagnostics, diagnostic scripting, and developer-facing automated diagnostic pipelines and interfaces.
This work addresses the challenge of diagnosing integration test failures, which is hindered by the massive volume, unstructured nature, and heterogeneity of logs, leading to inefficient root cause identification by developers. The authors propose a novel approach that deeply integrates large language models (LLMs) into Google Critique, an industrial-scale code review system, introducing a context-aware log comprehension and summarization method to automatically generate concise and accurate diagnostic insights. Evaluated on 71 real-world failure cases, the method achieves a 90.14% accuracy rate. Following deployment across 52,635 failed tests, only 5.8% of users reported the diagnostics as unhelpful, and the tool ranked 14th in helpfulness among 370 internal tools, demonstrating substantial improvements in both diagnostic efficiency and user experience.
This work proposes an intelligent agent-based diagnostic framework leveraging large language models (LLMs) to overcome the limitations of traditional root cause analysis methods, which rely on hard-coded rules, incur high maintenance costs, and are tightly coupled with infrastructure. By integrating a Model Context Protocol (MCP) and a constrained tool space, the framework enables agents to autonomously invoke tools for service querying, dependency retrieval, and multi-source data analysis, facilitating stepwise reasoning to pinpoint root causes. A structured investigation protocol ensures traceable and reproducible inference while maintaining robustness under incomplete or ambiguous information, effectively decoupling the model from underlying infrastructure. This approach lays the foundation for autonomous fault diagnosis and change impact assessment, paving the way for automated remediation and risk prediction, thereby significantly enhancing operational efficiency and system safety.
This work addresses the challenges in edge and embedded application development—namely, heterogeneous software stacks, multi-language runtimes, and difficult debugging—which lead to rigid deployment workflows and complex fault diagnosis. To overcome these limitations, the paper proposes a novel architecture enabling unified end-edge-cloud development. Its core components include a single programming language, a retargetable runtime system, a local recording and replay mechanism for distributed events, and a cross-platform deployment framework. This design breaks down traditional debugging barriers in edge–cloud collaborative development, facilitating seamless scalability, consistent testing, and flexible deployment across heterogeneous environments. Evaluation of the prototype system demonstrates that the proposed approach significantly simplifies deployment procedures and enhances fault diagnosis efficiency.
In complex equipment fault diagnosis, domain expertise is difficult to formalize, and manual fault tree construction is inefficient. Method: This paper proposes a knowledge graph–based approach for automated fault tree synthesis. It introduces a lightweight, semantically rich knowledge graph representation that enables semi-automatic extraction of failure logic relationships from unstructured documents (e.g., maintenance manuals) and structured/functional models. Leveraging hierarchical modeling and semantic reasoning, the method generates fault trees in a fully structured manner—without requiring historical fault data, relying solely on engineering knowledge. Contribution/Results: The synthesized fault trees are inherently interpretable and accurately capture system-level failure propagation paths. Experimental validation on the Lycoming O-320 aircraft engine demonstrates substantial improvements in diagnostic modeling efficiency and engineering applicability.
This work addresses the lack of automated, structured failure-recovery mechanisms in current software engineering agents, which struggle to translate heterogeneous runtime evidence into actionable repair guidance. The paper proposes PROBE, a novel framework that introduces a failure-anchored, structured recovery paradigm. PROBE employs a three-layer architecture—telemetry, diagnosis, and guidance gate—to decouple yet coordinate diagnosis and recovery, enabling non-intrusive integration. By integrating runtime telemetry, multi-signal diagnosis, and evidence-driven bounded guidance generation, PROBE constructs an end-to-end recovery pipeline. Evaluated on 257 unresolved cases, PROBE achieves a Top-1 diagnostic accuracy of 65.37% and a recovery success rate of 21.79%, significantly outperforming the strongest baseline. Its practical feasibility has been validated through deployment in Microsoft’s IcM system.
This work addresses the unreliability of large language model (LLM)-driven agent workflows, which stems from output nondeterminism, complex node dependencies, and tool heterogeneity, and proposes FlowFixer—a novel framework that introduces symbolic reasoning into automated workflow repair. FlowFixer models execution traces symbolically to generate behavioral specifications, enabling precise fault localization and root cause identification, and dynamically synthesizes targeted repair patches. To reduce verification overhead, it incorporates a multidimensional pre-evaluation mechanism. Experimental evaluation on Dify, Coze, and n8n platforms demonstrates that FlowFixer achieves a repair success rate of 71.3%, outperforming existing methods by 11.9%–27.6%, and improves root cause analysis accuracy by 15.3%–38.8%.
This study addresses the challenges of regression testing in remote and hybrid work environments, where communication, coordination, and quality assurance are increasingly complex. Through qualitative interviews with 20 software practitioners, complemented by process analysis, tool integration assessment, and coding of collaborative practices, the research systematically investigates the sociotechnical evolution of regression testing in distributed settings. Findings indicate that while core testing phases remain largely stable, teams increasingly rely on documentation, automation, and integrated toolchains to sustain effectiveness. Standardized reporting formats, shared repositories, and traceability mechanisms significantly mitigate collaboration barriers inherent in remote work. The study offers novel insights and practical guidance for ensuring software quality in geographically dispersed development contexts.
This study addresses the challenge of root-cause localization in automotive software testing, where high-dimensional sensor data generated during hardware-in-the-loop (HIL) simulations render traditional threshold-based methods ineffective. Existing data-driven approaches often require extensive labeled data and lack interpretability, failing to meet ISO 26262 traceability requirements. To overcome these limitations, the authors propose a two-stage diagnostic framework: first, safety requirements are automatically verified on a dSPACE real-time platform to filter anomalous test records; then, sliding windows of sensor signals are abstracted into statistical, relational, and contextual descriptors, which are fed as fixed prompts to an open-source large language model fine-tuned with 4-bit low-rank adaptation (LoRA). Evaluated on six fault types injected into a gasoline engine, the approach achieves 81.6% accuracy with a minimal 2B-parameter model—comparable to larger models—while operating entirely on a single consumer-grade GPU. This work pioneers the use of instruction-tuned large language models for sensor-level automotive fault diagnosis, demonstrating that diagnostic performance hinges more on task-specific adaptation convergence than on model scale, thereby achieving high accuracy, data efficiency, and explainable decision-making.
This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).