sdn traffic engineering

Designs, implements, and evaluates software-defined networking control-plane components (controllers, control applications, and rule-generation modules) that programmatically compute and install forwarding paths and policies to meet traffic-engineering objectives. Builds systems and experiments to enforce traffic policies, perform load balancing and failure recovery (including recreating failure-induced routing behaviors), and integrate with emulated topologies or testbeds to validate behavior.

sdntrafficengineering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.35
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$197K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Framework for Integrating Machine Learning Methods for Path-Aware Source Routing

Nov 17, 2024
AA
Anees Al-Najjar
🏛️ Oak Ridge National Laboratory | Federal Institute of Espírito Santo | Federal University of Espírito Santo | California Institute of Technology

To address the path optimization challenge for dynamic traffic engineering in software-defined networking (SDN), this paper proposes a real-time closed-loop control framework integrating deep reinforcement learning (DRL) with source routing. Methodologically, we design PolKA—a lightweight, P4-programmable source routing mechanism—and Hecate—a DRL-based system for real-time traffic analytics and path decision-making—achieving, for the first time, their coordinated closed-loop scheduling on a physical P4 testbed. Our key contribution is a data-plane-aware source routing integration paradigm that tightly couples path intelligence with programmable forwarding. Experimental results demonstrate a 37% reduction in end-to-end scheduling latency and a 52% decrease in link utilization variance, significantly enhancing network adaptability and operational controllability.

Machine LearningNetwork RoutingSoftware Defined Networking

This work addresses the vulnerability of data-driven security policies in software-defined networks (SDNs) to overreacting to anomalous traffic, which can lead to misclassification and degrade the performance of machine learning–based intrusion detection systems. To mitigate this issue, the authors propose Safeguard, a novel mechanism that introduces a set of allow rules derived from known benign traffic. These rules operate in conjunction with data-driven policies, enabling coordinated enforcement at the network edge to prevent unintended responses while simultaneously applying firewall rules against confirmed malicious traffic. By integrating this dual-layer approach, Safeguard effectively alleviates overblocking, significantly enhancing the robustness and accuracy of SDN security policies. Experimental evaluation through a prototype implementation demonstrates the efficacy of the proposed mechanism in dynamic SDN environments.

Data-driven PolicyIntrusion DetectionNetwork Security

This work addresses the challenges of in-band SDN control planes in resource-constrained wide-area telecommunication networks, including autonomous bootstrap, source routing, sub-50ms failure recovery, and multi-controller coordination. The authors propose Periplus, a system that embeds a forwarding graph—encoding both primary and per-hop backup paths—into L2/L3 packet headers. This design enables controller-independent local failover within 50 milliseconds and facilitates switch bootstrap with minimal flow table overhead. Notably, only two switches require initial configuration, eliminating network-wide flooding during provisioning. Flow table occupancy is decoupled from network size, scaling only at nodes that encode multipath information. Experimental evaluation using Ryu and Open vSwitch (augmented with Nicira extensions for NSH encapsulation) demonstrates Periplus’s capabilities in rapid recovery, scalable bootstrapping, and efficient resource utilization.

automatic bootstrappingfast failure recoveryin-band SDN

Model-Driven Rapid Prototyping for Control Algorithms with the GIPS Framework (System Description)

Mar 26, 2025
MK
Maximilian Kratz
🏛️ Technical University of Darmstadt

Software engineers face significant challenges—including difficulty in modeling, lengthy prototyping cycles, and high verification costs—when developing control algorithms for complex dynamic systems such as communication networks. To address these issues, we propose GIPS, the first model-driven engineering framework that tightly integrates graph-structured integer linear programming (ILP) modeling with automated code generation. Using the domain-specific language GIPSL, users declaratively specify constraints and optimization objectives; GIPS then automatically generates functionally complete, executable Java graph-optimization components. This enables end-to-end rapid prototyping—from high-level specifications to runtime deployment. We validate GIPS on a tree-structured peer-to-peer topology control scenario, demonstrating its correctness, efficiency, and scalability. The full implementation—including source code and a ready-to-run virtual machine demonstration environment—is open-sourced, confirming its practical deployability and engineering utility.

Automatically generates executable Java artifacts for graph optimizationDevelops GIPS framework for rapid prototyping of control algorithmsUses high-level language GIPSL to specify model optimization constraints

This paper addresses the challenge of detecting state-dependent performance issues (SPIs) in Software-Defined Networking (SDN) controllers—i.e., input sequences that drive the controller into anomalous states, causing severe performance degradation in subsequent operations. We propose the first dependency-aware, modular performance fuzzing methodology, integrating event-driven architecture modeling, static service dependency analysis, and state-sensitive coverage-guided grey-box fuzzing to systematically uncover SPIs across 157 network services in ONOS. Our approach identifies 10 previously unknown SPI vulnerabilities in ONOS, two of which have been confirmed to induce critical latency spikes or response blocking. Compared to existing performance fuzzing techniques, our method achieves significantly higher detection efficiency and uncovers deeper, more complex SPIs rooted in intricate service dependencies and state transitions.

Address large state space and complex architecture challengesDetect stateful performance issues in SDN controllersDevelop SPIDER framework for identifying network vulnerabilities

Latest Papers

What's happening recently
View more

This study addresses the coordination challenges in in-band SDN control plane deployments with multiple controllers, where controller discovery, state synchronization, and failure recovery must be achieved without expanding switch forwarding state. The authors propose a boundary-switch-based local forwarding graph mechanism that confines inter-domain routing information to boundary devices, preventing state propagation into intermediate domains. In-band control communication is realized using Open vSwitch’s Nicira extensions with NSH encapsulation, and neighbor discovery is accomplished via Controller Advertisement messages. The approach requires no switch firmware modifications and incurs flow table overhead independent of the number of controllers, maintaining constant space complexity. Experiments in a Mininet environment with 96 switches and 5 controllers demonstrate that internal switches exhibit fixed flow table occupancy, enabling network scalability to hundreds of nodes with controller discovery convergence times on the order of seconds.

controller failure recoveryforwarding statein-band SDN

This work addresses the limitations of centralized SDN in handling bursty traffic and the poor robustness of existing learning-based approaches under distribution shifts due to their reliance on offline training. The authors propose a hierarchical traffic control framework that enables online adaptation at edge nodes while adhering to global policy constraints. By employing policy envelopes to bound the action space per path, the framework supports real-time decisions on metering, queuing, and rerouting, ensuring local policies remain auditable and rollback-capable. Integrating centralized policy compilation with edge-based reinforcement learning, the approach demonstrates significant improvements in a 1024-host testbed: compared to Static ECMP, it achieves a 35.5% increase in core link utilization, a 34.3% reduction in P99 flow completion time for elephant flows, and lowers SLA violation rates from 18.2% to 6.8%, with each edge agent consuming less than 2% CPU and only 12 MB of memory.

Bursty TrafficDistribution ShiftSoftware Defined Networking

This study addresses the limitation of static load balancing in isolating backend servers that exhibit performance degradation—such as persistently returning HTTP 500 errors—without fully failing, which leads to elevated client-side error rates. The work presents the first systematic investigation into the feasibility of deploying open-source large language models (LLMs) within a real-world load-balancing control plane. Leveraging HAProxy and Prometheus telemetry updated every 10 seconds, the system dynamically quarantines faulty nodes via constrained API calls. Evaluations across 15 open-source LLMs spanning dense, mixture-of-experts (MoE), and sparse architectures reveal that models with approximately 3 billion active parameters represent a capability threshold for effective scheduling. Above this threshold, non-inference-based models reduce 5xx errors by 88%, albeit at the cost of 2.6–2.8× higher tail latency; enabling inference, however, increases token consumption tenfold and degrades control responsiveness.

backend degradationfault isolationHAProxy

This work addresses the challenge of automatically translating high-level service intents into effective Linux traffic control configurations, a task traditionally reliant on manual, low-level operations. The paper presents the first end-to-end framework that converts natural language or declarative intent specifications into standards-compliant Quality of Service (QoS) rules. The approach integrates queueing-theoretic semantic modeling, the LLaMA3 large language model, Active Queue Management (AQM)-guided prompting, and a rule-based validation mechanism to ensure correctness and compliance of the generated configurations. Experimental evaluation on 100 test intents demonstrates that LLaMA3 achieves a semantic similarity of 0.88 and a coverage of 0.87, outperforming baseline models by over 30%. Furthermore, AQM-guided prompting reduces output variability by a factor of three, significantly enhancing consistency and reliability.

intent-based networkingnetwork automationQuality of Service

Hot Scholars

KL

Kai Lei

Research Professor, Peking University, Shenzhen Graduate School
Future InternetData MiningBlockchain
ME

Muhammad Ejaz Ahmed

CSIRO's Data61
Securitydigital forensicsmalware detectionthreat hunting
RZ

Ruidong Zhu

Peking University
Machine Learning SystemsComputer SystemsDistributed Systems
XW

Xinming Wei

Peking University
Computer ArchitectureSecurityLLM
SO

Sofiane Ouni

INSAT, Carthage University (CRISALT LAB, ENSI)
Real Time Wireless Sensor networksIoT networks and applicationsIoT and BlockchainDistributed