LLM-Based Automated Diagnosis Of Integration Test Failures At Google

📅 2026-04-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of diagnosing integration test failures, which is hindered by the massive volume, unstructured nature, and heterogeneity of logs, leading to inefficient root cause identification by developers. The authors propose a novel approach that deeply integrates large language models (LLMs) into Google Critique, an industrial-scale code review system, introducing a context-aware log comprehension and summarization method to automatically generate concise and accurate diagnostic insights. Evaluated on 71 real-world failure cases, the method achieves a 90.14% accuracy rate. Following deployment across 52,635 failed tests, only 5.8% of users reported the diagnostics as unhelpful, and the tool ranked 14th in helpfulness among 370 internal tools, demonstrating substantial improvements in both diagnostic efficiency and user experience.

Technology Category

Natural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsMachine Learning: Large Multimodal Models (LMMs)Knowledge Representation and Reasoning: Diagnosis and Abductive Reasoning

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Integration testing is critical for the quality and reliability of complex software systems. However, diagnosing their failures presents significant challenges due to the massive volume, unstructured nature, and heterogeneity of logs they generate. These result in a high cognitive load, low signal-to-noise ratio, and make diagnosis difficult and time-consuming. Developers complain about these difficulties consistently and report spending substantially more time diagnosing integration test failures compared to unit test failures. To address these shortcomings, we introduce Auto-Diagnose, a novel diagnosis tool that leverages LLMs to help developers efficiently determine the root cause of integration test failures. Auto-Diagnose analyzes failure logs, produces concise summaries with the most relevant log lines, and is integrated into Critique, Google's internal code review system, providing contextual and in-time assistance. Based on our case studies, Auto-Diagnose is highly effective. A manual evaluation conducted on 71 real-world failures demonstrated 90.14% accuracy in diagnosing the root cause. Following its Google-wide deployment, Auto-Diagnose was used across 52, 635 distinct failing tests. User feedback indicated that the tool was deemed "Not helpful" in only 5.8% of cases, and it was ranked #14 in helpfulness among 370 tools that post findings in Critique. Finally, user interviews confirmed the perceived usefulness of Auto-Diagnose and positive reception of integrating automatic diagnostic assistance into existing workflows. We conclude that LLMs are highly successful in diagnosing integration test failures due to their capacity to process and summarize complex textual data. Integrating such AI-powered tooling automatically into developers' daily workflows is perceived positively, with the tool's accuracy remaining a critical factor in shaping developer perception and adoption.
Problem

Research questions and friction points this paper is trying to address.

integration test failures
log diagnosis
cognitive load
signal-to-noise ratio
software debugging
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-based diagnosis
integration test failure
log summarization
developer workflow integration
automated root cause analysis
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Celal Ziftci
Google
R
Ray Liu
Google
S
Spencer Greene
Google
L
Livio Dalloro
Google