Score
Designs, builds, or analyzes routing and traffic-management policies that integrate lightweight machine-learning models to infer runtime traffic conditions and produce traffic-aware routing decisions. These policies augment baseline shortest-path routing with learned predictions and drive adaptive scheduling and relocation decisions to reduce latency and improve relocation success rates.
For NP-hard routing optimization problems—such as the Traveling Salesman Problem (TSP) and Vehicle Routing Problem (VRP)—exact algorithms suffer from prohibitive computational complexity, while traditional heuristics lack optimality guarantees. This paper presents a systematic survey of machine learning (ML) applications in routing optimization and proposes a unified “constructive–improvement” taxonomy that integrates operations research models with deep learning (supervised and reinforcement learning), constructive heuristics, and local search techniques. Its key contributions are threefold: (i) the first structured, ML-driven classification framework for routing algorithms; (ii) a principled integration pathway bridging classical optimization and data-driven methods; and (iii) enhanced modeling and efficient solution capabilities for novel VRP variants. The work establishes a scalable theoretical foundation and practical paradigm for intelligent routing decision-making systems. (136 words)
To address the path optimization challenge for dynamic traffic engineering in software-defined networking (SDN), this paper proposes a real-time closed-loop control framework integrating deep reinforcement learning (DRL) with source routing. Methodologically, we design PolKA—a lightweight, P4-programmable source routing mechanism—and Hecate—a DRL-based system for real-time traffic analytics and path decision-making—achieving, for the first time, their coordinated closed-loop scheduling on a physical P4 testbed. Our key contribution is a data-plane-aware source routing integration paradigm that tightly couples path intelligence with programmable forwarding. Experimental results demonstrate a 37% reduction in end-to-end scheduling latency and a 52% decrease in link utilization variance, significantly enhancing network adaptability and operational controllability.
This work addresses the challenge that existing neural routing methods struggle to support near-real-time deployment in real networks due to telemetry data staleness caused by communication delays. The authors formulate telemetry-aware routing as a closed-loop control problem that explicitly integrates both communication and inference latencies. They propose LOGGIA, a scalable, localized graph neural routing framework that, for the first time, explicitly models these two delay components and combines pretraining with online policy reinforcement learning to enable each router to locally predict link weights in logarithmic space. Experiments on both synthetic and real-world topologies under unseen mixed TCP/UDP traffic demonstrate that LOGGIA significantly outperforms shortest-path baselines, while other neural approaches suffer notable performance degradation when realistic delays are introduced, thereby validating the efficacy of fully distributed deployment.
Traffic assignment is computationally expensive and impractical for real-time applications, especially on large-scale road networks. To address this, we propose an interpretable meta-model based on Message Passing Neural Networks (MPNNs), the first to align Graph Neural Network (GNN) architecture with the logic of Stochastic User Equilibrium (SUE) solving—directly mapping origin-destination (OD) demands to equilibrium flows without iterative simulation. The model takes a traffic graph as input and explicitly encodes path-choice behavior and flow allocation mechanisms, substantially improving out-of-distribution generalization. Experiments demonstrate that our approach reduces computational time by over 90% while preserving prediction accuracy, enabling real-time analysis on large-scale networks. Moreover, it exhibits strong robustness across distributionally shifted scenarios, overcoming the limited extrapolation capability typical of purely data-driven models.
To address path homogenization-induced congestion in dynamic multi-vehicle routing for urban environments—where shortest-path-first (SPF) algorithms often fail—the paper proposes a cooperative navigation framework based on multi-agent reinforcement learning. The method innovatively integrates graph attention networks (GATs) to jointly model local and neighborhood traffic states, introduces an adaptive navigation policy coupled with a hierarchical hub control mechanism, and employs a centralized-training-with-decentralized-execution (CTDE) architecture enhanced by Attentive Q-Mixing (A-QMIX) for traffic-aware global coordination. Evaluated on both synthetic and real-world road networks (Toronto and Manhattan), the approach reduces average travel time by up to 15.9% compared to SPF and state-of-the-art learning baselines, while maintaining 100% route success rate. It demonstrates superior scalability and congestion mitigation capability in large-scale urban transportation systems.
This paper addresses dynamic customer routing optimization in skill-based service queueing systems (e.g., cloud data centers), where static policies fail to adapt to workload fluctuations and heterogeneous agent skills. We propose a UCB-based reinforcement learning routing algorithm that jointly estimates environment dynamics, balances multiple objectives—namely, waiting time minimization and load balancing—and incorporates parameter sensitivity analysis. To accelerate convergence and enhance robustness, we innovatively integrate heuristic rules into the exploration mechanism. Extensive experiments driven by real-world operational data demonstrate that our algorithm significantly outperforms standard baselines in both efficiency and adaptability. It achieves rapid online learning and self-adjustment under varying traffic conditions, validating its practical feasibility and effectiveness for deployment in complex, large-scale service systems.
This work addresses the limitations of static penalty strategies in VLSI global routing, which struggle to adapt to complex congestion topologies and jointly optimize congestion, wirelength, and via count. To overcome these challenges, the paper proposes a dynamic multi-objective optimization framework that reformulates rip-up-and-reroute (R&R) as a dynamic system. The approach integrates SHAP-driven congestion decomposition, 3D Dijkstra maze routing, and an adaptive PathFinder algorithm, and—novelty introduced here—employs a large language model as a semantic policy optimizer to dynamically tune penalty parameters under knowledge graph constraints. Evaluated on the ISPD 2025 benchmarks, the method reduces MEMPOOL overflow by 98.6%, lowers ARIANE overflow to 146,109 (a 29.8× improvement over the state of the art), and achieves a penalty score of 0.0538, significantly outperforming the prior best result of 1.780.
This work addresses the absence of a unified framework for fairly comparing and efficiently deploying large language model (LLM) routing strategies under diverse query and budget constraints. We propose the first formalized LLM routing framework, modeling routing as a sequential decision process that integrates joint encoding of context and models, configurable scoring functions, flexible decision rules, and automated learning signals—supporting single-turn, multi-turn, and personalized routing. To facilitate research and development, we introduce xRouteBench, a multitask benchmark, and release LLMRouter, a modular infrastructure integrating over 16 routing methods. Experiments demonstrate that learned routers outperform the strongest fixed-model baseline by 14.6%, lightweight variants excel under stringent cost constraints, and user-conditioned routing significantly enhances personalization effectiveness.
This work addresses the challenge of efficient shortest-path planning in scenarios where real-world trajectory data are scarce, by leveraging simulators that exhibit systematic biases. The authors propose a graph Laplacian-regularized bias estimation method that integrates limited real observations, abundant synthetic data, and edge similarity structures within the road network to model smooth simulator-to-reality discrepancies. They establish theoretical guarantees on path suboptimality and devise an active learning strategy applicable even in the absence of initial real-world data. Through finite-sample error analysis and experiments on road networks across multiple cities, the approach demonstrates its ability to closely approximate optimal paths with only a small amount of real data, while providing computable performance certificates.
Telemetry-Aware routing promises to increase efficacy and responsiveness to traffic surges in computer networks. Recent research leverages Machine Learning to deal with the complex dependency between network state and routing, but sacrifices explainability of routing decisions due to the black-box nature of the proposed neural routing modules. We propose \emph{Placer}, a novel algorithm using Message Passing Networks to transform network states into latent node embeddings. These embeddings facilitate quick greedy next-hop routing without directly solving the all-pairs shortest paths problem, and let us visualize how certain network events shape routing decisions.