debug distributed systems

Designs, builds, and applies tools, instrumentation, and workflows to observe, reproduce, isolate, and fix faults across multi-node distributed systems and the hardware–software stack (including kernel and post‑silicon levels) in production and lab environments. Performs root‑cause analysis and troubleshooting using tracing, logging, breakpoints, hardware probes, deterministic replay, and related techniques, and develops debugging tools, processes, and test cases to validate fixes and prevent regressions.

debugdistributedsystems

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
2.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$196K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenges in edge and embedded application development—namely, heterogeneous software stacks, multi-language runtimes, and difficult debugging—which lead to rigid deployment workflows and complex fault diagnosis. To overcome these limitations, the paper proposes a novel architecture enabling unified end-edge-cloud development. Its core components include a single programming language, a retargetable runtime system, a local recording and replay mechanism for distributed events, and a cross-platform deployment framework. This design breaks down traditional debugging barriers in edge–cloud collaborative development, facilitating seamless scalability, consistent testing, and flexible deployment across heterogeneous environments. Evaluation of the prototype system demonstrates that the proposed approach significantly simplifies deployment procedures and enhances fault diagnosis efficiency.

cloud computingdistributed debuggingedge computing

This work addresses the inefficiencies and semantic inconsistencies arising from separately implementing driver and monitor programs in traditional hardware module testing. To overcome this, the authors propose a domain-specific language (DSL) tailored to hardware communication protocols, which enables the unified specification of both driver and monitor logic through an imperative syntax, thereby ensuring their semantic consistency for the first time. Building upon this DSL, they develop a prototype tool that leverages waveform parsing and transaction-level trace inference techniques to accurately reconstruct protocol-compliant transaction sequences from raw signal waveforms. Experimental results demonstrate that the approach significantly improves development efficiency, with further validation planned on real-world interconnect protocols such as Wishbone and AXI-Stream.

driverhardware communicationmonitor

Traditional interactive debuggers struggle to support distributed systems due to fragmented call stacks across process boundaries, difficulty in maintaining debugging state, and the risk of timeout cascades caused by global halting. This work proposes a novel approach enabling cross-process, source-level interactive debugging by embedding causal metadata into RPCs to reconstruct unified call stacks via distributed backtracing, designing an intent-preserving control plane that automatically coordinates dynamic breakpoints, and virtualizing logical clocks to eliminate temporal discontinuities. The system integrates lightweightly with frameworks such as gRPC, achieving a median cross-RPC backtrace latency of 30 ms, time skew under 5 ms, and throughput overhead of only 1–5% across 122 processes. User studies demonstrate a 100% success rate in fault localization with a median diagnosis time of approximately 8 minutes.

call stackdebugging statedistributed systems

This work addresses the inefficiency and heavy reliance on manual intervention in debugging counterexamples during formal connectivity verification. To this end, it introduces a novel graph-based perspective and proposes an automated root-cause analysis methodology. By integrating structural and functional dependency graphs generated by formal verification tools with counterexample reports, the approach leverages graph algorithms to classify verification failures and guide them into one of three targeted analysis workflows. This enables precise localization of fault points and provides either actionable repair suggestions or focused prompts for manual inspection. Evaluated on two industrial-scale SoCs, the method reduces debugging time by up to 80%, substantially enhancing debugging efficiency in complex scenarios and establishing the first systematic framework for automated debugging in connectivity verification.

counterexample debuggingdebug automationformal connectivity checking

Debugging embedded programs is notoriously challenging due to tight software-hardware coupling, and existing tools often rely on external hardware probes or serial logging, resulting in low efficiency. This work proposes Inline, a novel programming tool that, for the first time, enables real-time inline visualization of hardware logs directly within source code. It introduces a domain-specific expression language to support programmable manipulation of logs, allowing developers to intuitively trace execution flow and precisely localize faults. Seamlessly integrated into standard embedded development environments, Inline significantly lowers the barrier to effective debugging. A user study with twelve participants demonstrates marked improvements in both debugging efficiency and accuracy when using the tool.

embedded debugginghardware loginline visualization

Latest Papers

What's happening recently
View more

UniSage: A Unified and Post-Analysis-Aware Sampling for Microservices

Sep 30, 2025
ZZ
Zhouruixing Zhu
🏛️ The Chinese University of Hong Kong, Shenzhen | The Chinese University of Hong Kong

In modern distributed systems, massive trace and log data incur prohibitive storage overhead and impede fault diagnosis. Existing presampling methods often discard failure-relevant signals, compromising diagnostic transparency. This paper proposes UniSage—the first unified trace and log sampling framework tailored for microservices—adopting a *post-analysis–aware* paradigm: lightweight multimodal anomaly detection and root cause analysis (RCA) are first executed on the full data stream to generate service-level diagnostic insights that guide sampling decisions. UniSage innovatively integrates two complementary pillars: *analysis-guided sampling*, prioritizing critical failure signals, and *edge-case preservation*, ensuring rare but potentially diagnostic behaviors are retained. Experiments demonstrate that at a 2.5% sampling rate, UniSage captures 56.5% of critical traces and 96.25% of relevant logs, improves RCA accuracy@1 by 42.45%, and processes 10 minutes of data in under 5 seconds—substantially outperforming state-of-the-art approaches.

Improving accuracy of root cause analysis in distributed systemsPreventing loss of failure-related information in sampling approachesReducing storage overhead from growing microservices traces and logs

This work addresses the challenge of root cause diagnosis in distributed systems, where non-deterministic bugs often exhibit runtime symptoms far removed from their true origins, and existing tools struggle to capture complete cross-component causal evidence with low overhead. To this end, the paper presents Lumos, an online debugging framework that integrates static program analysis with dependency-guided instrumentation to selectively and lightweightly record program state histories relevant to observed anomalies. Lumos enables precise correlation between runtime symptoms and their root causes by reconstructing full causal chains on demand. It is the first application-layer approach to support low-overhead, automated, and on-demand fault tracing, requiring only a few observed bug instances to efficiently and accurately localize root causes.

distributed systemsnon-deterministic bugsonline debugging

This work addresses the challenge of constructing reproducible software stacks in high-performance computing (HPC) and AI convergence scenarios, where constraints such as lack of root privileges, network isolation, and heterogeneous language environments hinder conventional tooling. For the first time, it systematically applies Nix’s fully isolated build model together with its declarative flake configuration system to HPC-AI hybrid environments, enabling unified management of C/C++ and Python dependencies. By automatically generating Apptainer containers, the approach ensures consistency between local development and deployment on supercomputing systems. The method effectively resolves critical issues including dependency discovery, system library leakage, and cross-project composition, achieving highly reproducible deployments across non-root workstations and production clusters. It demonstrates clear advantages over traditional module systems, Conda environments, and manual containerization, while also highlighting current gaps in Nixpkgs’ coverage of machine learning packages.

dependency managementenvironment isolationHPC-AI software stack

This work addresses the challenge of debugging non-deterministic programs on microcontrollers, where sensor-driven inputs lead to irreproducible execution paths and existing techniques suffer from snapshot overhead, model dependency, and state explosion. The authors propose a trajectory-based “multi-verse” debugging approach that integrates concolic execution into this paradigm for the first time. By recording lightweight execution trajectories instead of full system snapshots, the method dynamically identifies critical inputs and prunes redundant paths, substantially mitigating state explosion while reducing memory and communication costs—making it suitable for resource-constrained environments. Implemented as a prototype atop the WARDuino WebAssembly virtual machine with a remote debugging architecture, the approach demonstrates significant reductions in state space and enhanced debugging efficiency and scalability in real-world scenarios compared to conventional solutions.

execution pathsmicrocontroller debuggingmultiverse debugging

This work addresses the lack of automated, structured failure-recovery mechanisms in current software engineering agents, which struggle to translate heterogeneous runtime evidence into actionable repair guidance. The paper proposes PROBE, a novel framework that introduces a failure-anchored, structured recovery paradigm. PROBE employs a three-layer architecture—telemetry, diagnosis, and guidance gate—to decouple yet coordinate diagnosis and recovery, enabling non-intrusive integration. By integrating runtime telemetry, multi-signal diagnosis, and evidence-driven bounded guidance generation, PROBE constructs an end-to-end recovery pipeline. Evaluated on 257 unresolved cases, PROBE achieves a Top-1 diagnostic accuracy of 65.37% and a recovery success rate of 21.79%, significantly outperforming the strongest baseline. Its practical feasibility has been validated through deployment in Microsoft’s IcM system.

failure diagnosispost-failure recoveryruntime telemetry

Hot Scholars

MP

Michael Pradel

Faculty, CISPA Helmholtz Center for Information Security • Professor, University of Stuttgart
Software EngineeringProgramming Languages
ZL

Zhou Liu

China Southern Power Grid/ Shenzhen Power Supply Co., Ltd.
Renewable Power IntegrationSmart gridPower system protectionDigital substation
WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc
YH

Yintong Huo

Singapore Management University
AI4SEAIOpsLog analysisMLLM for SE
RT

Radu Timofte

Humboldt Professor for AI and Computer Vision, University of Würzburg
Computer VisionMachine LearningAICompression