technical troubleshooting

Designs and implements diagnostic methods, monitoring/alerting configurations, runbooks, test harnesses, scripts and tools to detect, reproduce, isolate and remediate faults across hardware, infrastructure, distributed systems and production or field environments. Analyzes telemetry, logs, hardware diagnostics and customer reports to perform root-cause analysis, prioritize incidents, validate fixes, and produce corrective actions, mitigations, and post‑incident improvements for on‑call, customer‑facing, and on‑site operations.

technicaltroubleshooting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
2.38
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$187K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

ConfLogger: Enhance Systems' Configuration Diagnosability through Configuration Logging

Aug 28, 2025
SS
Shiwen Shan
🏛️ Sun Yat-sen University | Singapore Management University

Configurable systems frequently suffer from erroneous configurations and latent defects due to the complexity of their configuration spaces; existing diagnostic approaches predominantly focus on post-failure analysis and overlook the diagnostic potential embedded in software logs for configuration-related issues. This paper proposes a configuration-aware log enhancement method: it employs static taint analysis to identify propagation paths of critical configuration variables and leverages large language models (LLMs) to generate context-sensitive, semantically rich diagnostic log statements, while optimizing variable capture and code-context modeling. To our knowledge, this is the first automated configuration-log generation framework integrating static data-flow analysis with LLM-based semantic understanding. Evaluated across eight widely used systems, the method achieves 100% localization accuracy for 30 categories of silent misconfigurations, enables direct repair for 80% of such issues, improves diagnostic efficiency by 1.25×, and increases diagnostic accuracy by 251.4%.

Enhancing software configuration diagnosability through improved logging practicesGenerating diagnostic log statements using LLM-based code analysisIdentifying configuration-sensitive code segments via static taint analysis

This work proposes an intelligent agent-based diagnostic framework leveraging large language models (LLMs) to overcome the limitations of traditional root cause analysis methods, which rely on hard-coded rules, incur high maintenance costs, and are tightly coupled with infrastructure. By integrating a Model Context Protocol (MCP) and a constrained tool space, the framework enables agents to autonomously invoke tools for service querying, dependency retrieval, and multi-source data analysis, facilitating stepwise reasoning to pinpoint root causes. A structured investigation protocol ensures traceable and reproducible inference while maintaining robustness under incomplete or ambiguous information, effectively decoupling the model from underlying infrastructure. This approach lays the foundation for autonomous fault diagnosis and change impact assessment, paving the way for automated remediation and risk prediction, thereby significantly enhancing operational efficiency and system safety.

datacenter infrastructurefailure propagationincident diagnosis

From PREVENTion to REACTion: Enhancing Failure Resolution in Naval Systems

Aug 21, 2025
MT
Maria Teresa Rossi
🏛️ University of Milano -Bicocca

Naval systems frequently exhibit anomalous behaviors due to wear, misuse, or component failures—challenges that hinder timely detection and precise remediation. To address this, we propose a predictive-diagnostic closed-loop framework that tightly integrates the existing failure prediction system PREVENT with a newly designed responsive troubleshooting module, REACT. Methodologically, the framework synergizes multi-source time-series anomaly detection with domain-knowledge-driven fault-isolation process modeling, enabling end-to-end automation—from anomaly alerting and root-cause localization to actionable remediation recommendations. Evaluated on operational shipboard systems deployed by Fincantieri, the framework reduces mean time to fault localization by 42%, significantly improves operational response efficiency, and demonstrates strong generalizability across diverse industrial domains.

Enhancing failure detection and resolution in naval systemsExtending predictive maintenance to industrial productsIntegrating anomaly detection with troubleshooting procedures

This work addresses the challenges of root cause diagnosis in large-scale microservice systems, where existing approaches are hindered by massive log volumes, limited LLM context windows, and insufficient semantic reasoning and interpretability. The authors propose a neuro-symbolic hybrid method that emulates Site Reliability Engineers’ manual troubleshooting process through a six-stage pipeline for log sampling, template clustering, and anomaly ranking, producing a concise evidence package for LLM-based root cause inference. This approach compresses raw logs by 1,000–7,000× while preserving critical failure signals and provides auditable log templates and statistical evidence, substantially enhancing interpretability and practicality. Evaluated on 11 real-world incidents, the method achieves an MRR of 0.790 and ranks the correct root cause within the top three candidates in over 90% of cases within one minute, earning strong endorsement from operations teams.

incident diagnosislarge-scale systemslog analysis

A Two-Staged LLM-Based Framework for CI/CD Failure Detection and Remediation with Industrial Validation

Jun 04, 2025
WX
Weiyuan Xu
🏛️ East China Normal University | ByteDance

CI/CD pipeline failure diagnosis and repair have long suffered from high complexity and low automation. This paper introduces LogSage—the first end-to-end, LLM-driven framework for root-cause analysis (RCA) and automated repair. It features a novel two-stage LLM architecture: (1) Stage I employs intelligent log preprocessing to precisely localize failures; (2) Stage II integrates retrieval-augmented generation (RAG) with tool calling to generate executable, validated fixes. LogSage is the first industrial-grade solution validated on over one million production CI/CD pipelines. It achieves 98% RCA accuracy—12 percentage points higher than state-of-the-art baselines—and end-to-end repair accuracy exceeding 88%. Deployed at scale, it supported 1.07 million CI/CD executions in its first year, processing over 3,000 tasks daily.

Detect and fix CI/CD pipeline failures automaticallyImprove precision in root cause analysis of logsIntegrate historical solutions for actionable fixes

Latest Papers

What's happening recently
View more

This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).

Agentic SystemsMonitoringStructural Defects

This work addresses the pervasive issue of redundant and isolated messages in system logs, which hinder downstream tasks such as model reasoning and anomaly detection. To tackle this challenge, the authors propose LogPurifier—the first task-agnostic log cleansing framework—that systematically purifies logs by extracting log templates and modeling their dependencies to accurately identify and remove messages irrelevant to system functional behavior. By doing so, LogPurifier enables effective log sanitization applicable across diverse analytical scenarios. Experimental results demonstrate that LogPurifier substantially improves both accuracy and efficiency in various downstream tasks, thereby validating its effectiveness and generalizability.

downstream tasksirrelevant messageslog analysis

This work addresses the lack of automated, structured failure-recovery mechanisms in current software engineering agents, which struggle to translate heterogeneous runtime evidence into actionable repair guidance. The paper proposes PROBE, a novel framework that introduces a failure-anchored, structured recovery paradigm. PROBE employs a three-layer architecture—telemetry, diagnosis, and guidance gate—to decouple yet coordinate diagnosis and recovery, enabling non-intrusive integration. By integrating runtime telemetry, multi-signal diagnosis, and evidence-driven bounded guidance generation, PROBE constructs an end-to-end recovery pipeline. Evaluated on 257 unresolved cases, PROBE achieves a Top-1 diagnostic accuracy of 65.37% and a recovery success rate of 21.79%, significantly outperforming the strongest baseline. Its practical feasibility has been validated through deployment in Microsoft’s IcM system.

failure diagnosispost-failure recoveryruntime telemetry

This work addresses the challenge of effectively analyzing massive, heterogeneous high-performance computing (HPC) logs, which hinders fault diagnosis and performance optimization. The authors propose a scalable log analysis workflow that uniquely integrates frequent pattern mining based on finite-state automata with job-level log correlation. By leveraging the Aho–Corasick automaton for efficient pattern storage and matching, and incorporating system hierarchy and message priority information, the approach enables automated detection and clustering of errors and anomalous events. Experiments on an exascale-class supercomputing system demonstrate that the method accurately identifies characteristic error sequences, reveals distinct failure patterns across different applications, and supports real-time, interpretable monitoring to enhance system resilience.

anomaly detectionHPC logslog analysis

Hot Scholars

YL

Yang Liu

Peking University
Computer VisionMulti-modal Learning
YW

Yawen Wang

The University of Texas at Arlington
Gear DynamicsNoise and Vibration
EB

Emad Barsoum

AMD, Columbia University
Generative AIFoundation ModelsAgentic AIComputer Vision
FS

Federica Sarro

Professor, University College London
AI EngineeringSBSEAutomated Software EngineeringEmpirical Software Engineering
XW

Xiaolong Wu

Georgia Institute of Technology
SLAMLocalizationRobotics