network diagnostics

Designs and builds tools, procedures, and analyses to capture, inspect, and interpret network traffic and protocol exchanges in order to detect, localize, and diagnose faults and performance problems across wired and wireless systems. Work covers packet-capture and protocol analysis, traffic and congestion analysis, cross‑layer and production‑traffic investigation, and debugging of interactions between network, system, and hardware components.

networkdiagnostics

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.49
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Compact Data Structures for Network Telemetry

Nov 05, 2023
SL
Shir Landau Feibish
🏛️ The Open University of Israel | University of Maryland | Princeton University

Conventional network telemetry frameworks struggle to support fine-grained traffic measurement, performance diagnostics, and attack detection under stringent memory and computational constraints of high-speed network devices. Method: This paper proposes a lightweight, real-time online telemetry framework that systematically integrates compact data structures—including Bloom filter variants, Count-Min Sketch, and HyperLogLog—with streaming algorithms, hierarchical sampling, and P4-programmable data-plane co-design to comply with hardware limitations. Contribution/Results: Evaluated at line rate exceeding 100 Gbps, the framework reduces memory footprint by over 60% compared to state-of-the-art approaches while maintaining sub-1% flow frequency estimation error. It achieves an optimal trade-off among accuracy, throughput, and resource overhead, thereby significantly enhancing the feasibility and practicality of telemetry in high-bandwidth environments.

Compact data structures for traffic analysisHigh-speed network device limitationsTrade-offs between accuracy and overhead

Verification and Attack Synthesis for Network Protocols

Nov 02, 2025
MV
Max von Hippel
🏛️ Northeastern University

Ensuring functional correctness and performance resilience of network protocols under component failures and adversarial attacks remains a significant challenge. Method: This paper proposes a synergistic analysis framework integrating formal verification with attack synthesis. It models protocol behavior using a formal specification language and employs logical predicates, trace analysis, and model checking to achieve closed-loop verification—simultaneously establishing correctness guarantees and automatically generating realistic attack scenarios. Contribution/Results: Diverging from conventional unidirectional verification, our approach innovatively embeds attack-path generation directly into the verification workflow, enabling reproducible and interpretable failure attribution. Experimental evaluation across multiple mainstream network protocols demonstrates substantial improvements in vulnerability detection rates and attack-surface characterization accuracy. The results validate the feasibility and practicality of formal methods for deep, security-critical analysis of complex network protocols.

Applying formal methods to analyze protocols under normal and attack conditionsSynthesizing attacks that prevent protocol requirement achievementVerifying network protocol functionality and performance requirements

Model-Based Diagnosis: Automating End-to-End Diagnosis of Network Failures

Jun 29, 2025
CW
Changrong Wu
🏛️ Microsoft | UCLA | Cisco Systems | Alibaba Cloud

Rapid root-cause diagnosis of enterprise network failures has long been hindered by fragmented diagnostic scope—limited to either the data or control plane—and heavy reliance on manual analysis. Method: This paper proposes a model-based network diagnosis paradigm that, for the first time, systematically unifies hardware, firmware, and software-layer faults, jointly modeling both the data plane and distributed control plane. Leveraging network verification techniques, it constructs end-to-end forwarding and routing models, enabling automated root-cause inference from user-level symptoms within P4 switch and distributed routing software simulation environments. Contribution/Results: Evaluation demonstrates 100% diagnostic accuracy in simulation. On 33 real-world failure cases from a major cloud provider, the approach localizes root causes for 30 cases within seconds—achieving over two orders-of-magnitude speedup versus manual analysis—and significantly advances automation in enterprise network operations.

Accelerates diagnosis compared to human operatorsAutomates end-to-end diagnosis of network failuresCovers data plane and control plane faults systematically

This work addresses the limitations of traditional passive network measurement, which primarily focuses on inbound traffic and struggles to detect stealthy internal anomalies. The paper presents the first systematic approach that leverages erroneous outbound traffic—such as unanswered requests and ICMP error messages—as a lightweight yet highly informative data source. By conducting large-scale passive monitoring and correlation analysis, the method effectively identifies misconfigurations, deprecated services, and potentially compromised hosts within internal networks. Deployed in large-scale operational environments, this technique has uncovered a variety of previously undetected internal anomalies, substantially enhancing visibility into and detection capabilities for internal threats.

erroneous outbound trafficICMP errorsinternal anomalies

Leveraging Large Language Models to Contextualize Network Measurements

May 25, 2025
RB
Roman Beltiukov
🏛️ UC Santa Barbara

Non-technical users often misinterpret network measurement data (e.g., latency, packet loss, throughput), leading to erroneous conclusions. To address this, we propose the first systematic framework leveraging large language models (LLMs) for semantic interpretation of network measurements. Our method integrates historical measurement data with context-aware prompt engineering to automatically translate raw metrics into natural-language performance explanations, enabling scenario-adaptive and personalized feedback. Key contributions include: (1) introducing the first context-aware LLM reasoning paradigm specifically designed for network measurements, overcoming the limitations of conventional threshold-based alerting; and (2) significantly improving non-experts’ comprehension accuracy of critical metrics—experiments show a 37.2% average improvement—while delivering real-time, interpretable, and low-barrier diagnostic support.

Automating insights from low-level network metric dataEnhancing accessibility of network performance diagnosticsInterpreting network measurements for non-technical users

Latest Papers

What's happening recently
View more

This study addresses the challenge of isolating performance anomalies in the middle-mile segment of Internet paths—such as topology errors, suboptimal routing policies, and interconnection congestion—from end-host effects. Leveraging Measurement Lab (M-Lab) data, the authors employ a natural experiment design: users from the same access ISP connect to multiple geographically proximate M-Lab servers, enabling an A/B comparison that effectively controls for client-side, access-network, and temporal variability. This approach, applied at scale for the first time, uncovers previously masked middle-mile anomalies and enables joint detection of topological, routing, and congestion issues. Using a sparse multidimensional histogram method on BigQuery, the system computes Kolmogorov–Smirnov distances and geometric mean throughput ratios in a single pass over millions of samples, efficiently identifying bandwidth bottlenecks, traffic shaping, and suboptimal routes. Results are made publicly accessible through a metropolitan-level real-time dashboard supporting fine-grained analysis.

anomalous topologycongested interconnectionsInternet-scale measurement

This work addresses the critical vulnerability of Software-Defined Networking (SDN) controllers in Wide Area Networks (WANs) to severe outages caused by erroneous inputs, such as those stemming from control-plane bugs. To mitigate this risk, the authors introduce input validation as a dedicated defense layer within the WAN control plane, deploying a lightweight validation mechanism ahead of the controller to detect and block invalid inputs in real time. The system employs a shadow deployment architecture that combines simulation with live production data, exhibiting strong robustness against noisy, missing, or corrupted telemetry. During a four-week production deployment, it accurately captured the sole instance of invalid input with zero false positives. Simulations further demonstrate 100% detection of traffic anomalies as small as 5% and sustained zero false positives even under up to 30% telemetry corruption.

input validationinvalid inputsnetwork outages

This work proposes a novel interactive analysis system centered on three-dimensional network topology to overcome the limitations of traditional PCAP analysis tools, which present data as linear lists and fail to reveal underlying communication structures. The system maps hosts, sessions, and protocols to nodes, edges, and visual clusters, respectively, and enables bidirectional synchronized filtering with the packet list. By adopting 3D space as the default view—implemented using Three.js—it intuitively encodes key features such as communication density, clustering structure, host centrality, and traffic volume through depth perception. Supporting parsing of PCAP/PCAPNG formats and decoding of over 90 protocols, the approach significantly enhances the observability of structural patterns in network traffic, facilitating efficient identification of anomalous communications, critical nodes, and protocol distributions.

interactive visualizationnetwork topologypacket analysis

Existing evaluation methodologies struggle to emulate the multi-vendor, multi-protocol, and partially observable environments characteristic of real-world telecommunications networks, thereby failing to effectively assess AI agents’ capabilities in fault diagnosis and path restoration. To address this gap, this work proposes CTBench—the first public benchmark tailored for telecom operations—focusing on root cause analysis and path restoration tasks. Developed with expert-crafted simulation scenarios and annotated gold-standard evidence chains, CTBench introduces an expert-aligned evaluation metric that accounts for both final answers and diagnostic reasoning, emphasizing interpretability and evidential support, while also providing fine-grained metadata. Experimental results demonstrate that state-of-the-art agents perform reasonably well in path restoration but exhibit significant deficiencies in complex root cause analysis—particularly involving interface states, link-layer issues, and service management failures—and often lack sufficient diagnostic evidence to substantiate their conclusions.

AI agent evaluationpath restorationroot cause analysis

This work addresses the limitations of existing 802.11 packet capture diagnostics, which either rely heavily on expert knowledge—resulting in low efficiency and poor scalability—or employ large language models (LLMs) prone to hallucination, unreliable confidence estimates, and evaluation biases due to circular reasoning. To overcome these challenges, the authors propose PROBE, an evidence-driven, multi-stage diagnostic pipeline that integrates deterministic frame-level PCAP textualization, multi-model and multi-candidate ensembling, progressive masking, and a composite reliability scoring mechanism that operates without LLM-based self-evaluation. Evaluated on 87 enterprise Wi-Fi packet captures, PROBE achieves an evidence F1 score of 0.957 and a 96% automatic acceptance rate, with worst-case F1 exceeding 0.70—significantly outperforming both expert baselines (0.871) and naive ensemble approaches (0.842).

802.11 packet capture diagnosisevaluation biasexpert scalability

Hot Scholars

MT

Mina Tahmasbi Arashloo

University of Waterloo
Networked SystemsSoftware Defined NetworksProgrammable Networks
XZ

Xinggong Zhang

Peking University
AI-driven Multimedia NetworkingVideo CommunicationTransport Protocol
RB

Ryan Beckett

Microsoft Research
networkingprogramming languagesformal methodsverification
RB

Raouf Boutaba

Professor of Computer Science, University of Waterloo
Network ManagementNetwork VirtualizationCloud Resource Management
JH

John Heidemann

University of Southern California / Information Sciences Institute
networkingInternet mappingfile systems