manage mobility

Designs, builds, and evaluates systems and mechanisms that enable moving compute instances, network sessions, or device endpoints between hosts or network points while preserving state and session continuity. This includes atomic workload migration, updating traffic-routing and policy state, coordinating cross-node handovers, and maintaining low-latency connectivity during transitions.

managemobility

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Seamless Transitions: A Comprehensive Review of Live Migration Technologies

Dec 03, 2025
SA
Sima Attar-Khorasani
🏛️ TUD Dresden University of Technology | Center for Scalable Data Analytics and Artificial Intelligence

Existing surveys on live migration predominantly emphasize theoretical mechanisms while neglecting practical deployment constraints and technical adaptability. This paper systematically compares live migration techniques for containers and virtual machines across three dimensions—migration mechanisms, migration units, and infrastructure characteristics—to analyze performance, overhead, and compatibility trade-offs. It introduces, for the first time, an integrated evaluation framework that jointly considers migration objectives (e.g., cloud-edge coordination), operational constraints (e.g., resource and network limitations), and adoption heterogeneity, thereby enabling scenario-aware technology assessment and evolutionary guidance. Through multidimensional empirical analysis—including pre-copy/post-copy strategies, dirty-page tracking, CRIU-based checkpointing, KVM/QEMU live migration, and CRI-O hot migration—the study identifies five pervasive challenges. The findings provide actionable, deployment-oriented insights for technology selection and optimization in elastic scheduling and other production-critical scenarios.

Addresses gaps in existing reviews of live migration technologiesAnalyzes migration techniques, units, and infrastructure characteristics comprehensivelyExamines challenges and adoption disparities in container and VM migration

This paper identifies a critical gap in high-performance data transfer research: an overemphasis on network bandwidth while neglecting end-to-end bottlenecks—including latency, TCP congestion control, host CPU limitations, and virtualization—leading to severe discrepancies between benchmark results and real-world production performance. To address this, the authors propose a hardware–software co-design paradigm and develop a latency-programmable testbed. Leveraging high-fidelity wide-area network (WAN) modeling and cross-continental 100 Gbps measurements (Switzerland–California), they systematically isolate key constraints at the network edge and host side. Results demonstrate that primary bottlenecks reside predominantly at the network edge—not the core—and that stable, predictable data movement is achieved across 1–100+ Gbps. This significantly enhances performance fidelity in complex, heterogeneous environments.

Examines host-side factors like CPU and virtualization impacting workflows.Investigates bottlenecks beyond network bandwidth in data movement.Proposes holistic hardware-software co-design for consistent performance.

Building State Machine Replication Using Practical Network Synchrony

Jul 17, 2025
YW
Yiliang Wan
🏛️ National University of Singapore | Agency for Science, Technology and Research (A*STAR) | ETH Zurich

Traditional state machine replication (SMR) suffers from performance limitations in partially or fully asynchronous networks. This paper proposes Chora, an SMR protocol designed to exploit the strong synchronization characteristics of modern datacenter networks, enabling low-overhead, highly parallel replication. Chora employs kernel-bypass networking, a multi-threaded architecture, and relaxed round boundary control—achieving tightly bounded 2-μs rounds—to realize pipelined, multi-instance parallel replication. Crucially, it permits concurrent, coordination-free proposal generation across replicas. Its core innovation lies in directly translating network-level synchronization into protocol-level efficiency gains, thereby eliminating the serialization and coordination overhead inherent in conventional consensus protocols. Experimental results demonstrate that Chora achieves 255% and 109% higher throughput than the best-performing single-leader and multi-leader SMR protocols, respectively, significantly surpassing existing SMR performance bottlenecks.

Designing state machine replication with strong network synchronyImproving throughput via parallel replication without extra coordinationLeveraging modern data centers for synchronous distributed protocols

Transparent and Efficient Live Migration across Heterogeneous Hosts with Wharf

Oct 21, 2024
YY
Yiwei Yang
🏛️ UC Santa Cruz | ShanghaiTech University | Santa Clara University

LLM agents face significant challenges in cross-heterogeneous-host migration (e.g., x86/ARM, Linux/FreeBSD), including data confidentiality, low-latency responsiveness, high availability, and output integrity. Method: This paper proposes Vessel, a lightweight WebAssembly-based containerization framework. It introduces WASI as a unified abstraction for process state and OS interfaces, and designs the Dock runtime to enable dynamic migration-safety-point detection, cross-OS coordination, and latency-aware triggering—achieving transparent, real-time migration without code modification or service restart. Contribution/Results: Experiments demonstrate a 57% reduction in migration pause time while ensuring sensitive data remains within trusted boundaries and critical outputs are integrity-protected. Vessel natively supports C/Rust binaries, enabling load balancing, hot updates, and fault tolerance. It establishes a novel paradigm for trustworthy AI agent deployment across diverse infrastructure.

Minimizing response latency and ensuring output safety for critical applicationsProtecting sensitive user data while maintaining availability during failuresSecurely executing and migrating AI agents across heterogeneous environments

This work addresses the challenge of supporting voluntary GPU sharing in campus environments, where resources are subject to revocation and existing migration mechanisms—assuming static failures and unlimited transfer windows—are inadequate. The authors propose a network-layer live migration protocol that treats resource reclamation as a first-class contract, explicitly modeling provider-initiated revocation as a core constraint. Their approach co-optimizes reclamation-aware checkpoint scheduling, volatility-aware destination node selection, and sub-millisecond deadline-aware traffic control leveraging TC BPF. Evaluated over two months on a 54-node heterogeneous testbed, the system reduces job loss by 66% compared to Slurm preemption with requeueing and by 38% against pipelined redundant checkpointing, while cutting downtime by 38% and incurring less than 3% performance degradation on background scientific workloads.

GPU sharingheterogeneous infrastructurenetwork migration

Latest Papers

What's happening recently
View more

This study addresses cascading failures caused by shared firmware in commercial multi-host network interface cards (NICs) and the operational challenges of hyperscale deployments by proposing fbnic, a system-level solution. Architecturally, it introduces physical isolation and driver-priority mechanisms, combined with sub-sled-granularity firmware upgrade orchestration and slice-level fault containment. Operationally, it establishes a hardware-in-the-loop (HIL) continuous integration pipeline alongside cross-layer fault attribution monitoring and an automated remediation toolchain. Deployed across hundreds of thousands of hosts, fbnic reduces unplanned unavailability by 12×, shortens mean time to repair by 37%, and decreases hardware replacement rates by 2.3×, demonstrating robust stability at scale.

custom hardwarehyperscalerisolation failures

Traditional flow-level load balancing often suffers from hash collisions and tail latency due to its lack of link-state awareness, while existing packet-level approaches are hindered by stale state information, high hardware overhead, and poor adaptability to heterogeneous links. This work proposes the Probabilistic State Proportional (PSP) scheduling algorithm, which innovatively combines discrete state modeling based on bandwidth intervals with localized probabilistic mapping to achieve low-overhead, highly stable packet-level load balancing without requiring real-time global information. Simulations demonstrate that PSP significantly outperforms Join-the-Shortest-Queue (JSQ) and Random across diverse network scales and interference scenarios, achieving lower packet loss rates and 99th-percentile buffer occupancy, better scalability, and performance comparable to Top-k at substantially reduced hardware cost.

bandwidth asymmetrydata center networksload balancing

Hot Scholars

HD

Hans D. Schotten

Univ. of Kaiserslautern, RPTU Kaiserslautern, DFKI GmbH
Mobile and wireless communicationsindustrial radioindustrial internetsecurity
KH

Kaibin Huang

Professor and Dept.Head, University of Hong Kong; NAI Fellow; IEEE Fellow; Highly Cited Researcher
Machine LearningMobile Edge ComputingWireless CommunicationsWireless Power Transfer
HY

Halim Yanikomeroglu

Chancellor’s Professor, Systems and Computer Engineering, Carleton University, Canada
6GWireless CommunicationsNon-Terrestrial Networks5G
MS

Min Sheng

Xidian University
Mobile communication systemsAd hoc networksCognitive wireless networksHeterogeneous networks
GG

Giovanni Geraci

Nokia | Universitat Pompeu Fabra
AI/ML6GWi-FiWireless Communications