Score
Designs and implements mechanisms and protocols to transfer a running process's execution state (CPU registers, memory pages, open descriptors, and scheduling metadata) across hosts or execution contexts while preserving live execution slices. Builds, integrates, and evaluates migration components with operating-system schedulers and resource managers to ensure correctness, low-latency handoff, minimal disruption (e.g., avoiding tenant preemption mid-slice), and acceptable performance and consistency.
Existing surveys on live migration predominantly emphasize theoretical mechanisms while neglecting practical deployment constraints and technical adaptability. This paper systematically compares live migration techniques for containers and virtual machines across three dimensions—migration mechanisms, migration units, and infrastructure characteristics—to analyze performance, overhead, and compatibility trade-offs. It introduces, for the first time, an integrated evaluation framework that jointly considers migration objectives (e.g., cloud-edge coordination), operational constraints (e.g., resource and network limitations), and adoption heterogeneity, thereby enabling scenario-aware technology assessment and evolutionary guidance. Through multidimensional empirical analysis—including pre-copy/post-copy strategies, dirty-page tracking, CRIU-based checkpointing, KVM/QEMU live migration, and CRI-O hot migration—the study identifies five pervasive challenges. The findings provide actionable, deployment-oriented insights for technology selection and optimization in elastic scheduling and other production-critical scenarios.
LLM agents face significant challenges in cross-heterogeneous-host migration (e.g., x86/ARM, Linux/FreeBSD), including data confidentiality, low-latency responsiveness, high availability, and output integrity. Method: This paper proposes Vessel, a lightweight WebAssembly-based containerization framework. It introduces WASI as a unified abstraction for process state and OS interfaces, and designs the Dock runtime to enable dynamic migration-safety-point detection, cross-OS coordination, and latency-aware triggering—achieving transparent, real-time migration without code modification or service restart. Contribution/Results: Experiments demonstrate a 57% reduction in migration pause time while ensuring sensitive data remains within trusted boundaries and critical outputs are integrity-protected. Vessel natively supports C/Rust binaries, enabling load balancing, hot updates, and fault tolerance. It establishes a novel paradigm for trustworthy AI agent deployment across diverse infrastructure.
This study addresses the issues of residual stale outputs and loss of valid state during task revision in LLM agents by proposing a versioned execution control plane. The proposed method transforms abort-and-restart operations into seamless version transitions, decouples resource scheduling from release permissions, and unifies fast-fail with selective retention mechanisms to support the inheritance of completed KV cache states. A vLLM-based system implementation encompasses GPU execution, hierarchical recovery, and multi-tenant serving. Experimental results demonstrate that this approach reduces the median time-to-first-token latency by 17.1% while eliminating obsolete outputs and effectively preserving execution progress across repeated revisions.
本文通过对比SPEC CPU2026基准测试与其上游开源版本在单副本和多副本运行场景下的性能,量化分析了两者之间的'保真度差距',验证了SPEC方法的有效性。
本文提出XMPIaaS系统,通过协作式进程迁移解决MPI应用在云环境中因资源临时性和动态定价导致的执行问题。
To address the inefficiency caused by blind page migration in multi-tenant hierarchical memory systems, this paper proposes an adaptive page migration decision framework. The method introduces a novel, per-page “ping-pong” behavior–based dynamic migration-friendliness detection mechanism, enabling process-granularity control over migration initiation and suspension. It jointly models page hotness and migration patterns to support low-overhead, fine-grained runtime adaptation to access pattern changes. Implemented at the Linux kernel level and co-optimized with CXL hardware, the framework leverages hardware-assisted memory semantics for efficient decision-making. Evaluation on commercial CXL platforms demonstrates substantial reduction in spurious migrations: average memory latency decreases by 12.7% and bandwidth utilization improves by 19.3% across both single- and multi-tenant workloads.
This study addresses the reliance on manual configuration when deploying high-performance computing (HPC) workflows across heterogeneous sites by proposing an evidence-based method for the automated generation of structured site configurations. The approach leverages constrained language model agents to extract relevant information from documentation, combined with automated node profiling and pilot job execution for rule validation, thereby bridging the critical knowledge gap concerning site-specific usage practices in workflow portability. Experimental evaluations conducted at Purdue, TACC, and Notre Dame demonstrate that the proposed method successfully constructs accurate configuration files, enabling automated pre-flight verification of real-world workflows while providing interpretable failure diagnostics.
为解决固件重宿中外围设备建模不准确的问题,提出反应式外围建模(RPM)方法,通过事件-条件-动作语义忠实再现复杂外设行为。
This study addresses the GPU resource contention between online agent self-evolution and real-time serving, which causes delayed returns and high recovery costs. We propose LearnSched, a scheduler that incorporates capability reuse windows and computation regression times into a unified investment model to dynamically determine evolution actions. Methodologically, we construct a state-aware scheduling framework integrating candidate progress, checkpoint overhead, and recovery paths for global optimization, alongside a one-step counterfactual reasoning algorithm to select optimal execution strategies. Experimental results demonstrate that our approach significantly improves net system value under cold-recovery scenarios, validating that preserving recoverable states effectively shortens capacity regression cycles and reduces scheduling complexity.
本文通过将Paxos协议的伪代码转化为可执行的DistAlgo语言,解决了分布式系统中复制与共识协议的理解和验证问题。
本文提出一种运行时独立架构,以解决长期存在的代理在更换模型、编排框架、交互会话和主机服务器时保持同一身份、记忆和代码的问题。