🤖 AI Summary
This study addresses the invalidation of scheduling decisions caused by dynamic state changes in the computing continuum by proposing a digital twin-based scheduling framework. Through a synchronized physical-digital dual execution mechanism, the framework evaluates candidate decisions and verifies their feasibility, enabling for the first time the direct exploration of alternatives from the running state while supporting decision rehearsal and real-time state validation. The proposed method integrates DAG workload simulation with proximal policy optimization (PPO) reinforcement learning for joint optimization. Experimental results demonstrate that the system achieves a prediction error of only 0.485 seconds, reduces decision rejection latency to as low as 2 milliseconds, and improves the performance of weaker models by up to 29.1%.
📝 Abstract
Computing-continuum applications distribute work across devices, edge systems, fog resources, and clouds. While a placement, scheduling, or recovery decision is being made, resource availability, network conditions, and application progress may change, so the decision can be invalid by the time it is executed. Existing runtimes enact predetermined decisions, whereas simulation tools compare alternatives in preconfigured environments; neither explores alternative outcomes directly from the current state of a running application. We propose a Digital Twin framework, called Darpan, that supports both physical and digital execution: the physical side runs real applications and continuously observes their runtime state, while the digital side maintains a continuously updated, executable virtual counterpart based on these physical observations. Darpan evaluates candidate decisions independently from the same starting point and returns the selected decision to the physical side, where its feasibility is validated against the latest physical state before actual execution. Across real Directed Acyclic Graph (DAG) workloads, Darpan predicts physical response time with a mean absolute error of 0.485 s and retains 93.1% scale-out efficiency at 40 nodes. In the state-change trials, Darpan takes about 2 ms on average from state capture to rejection of an invalidated decision, stopping the request before data transfer and avoiding unnecessary physical execution overhead. Under the same physical-DAG budget, Darpan-generated experience enables a weaker Proximal Policy Optimization (PPO) scheduler to outperform two state-of-the-art Deep Reinforcement Learning (DRL) schedulers by 29.1% and 20.2%.