Schedule Repair for DAG Workflows under Link Disruptions
This study addresses scheduling failures in DAG workflows for IoT systems under adversarial conditions, where link disruptions compromise execution. We propose a multi-level repair strategy spectrum accounting for decision latency, spanning from passive waiting to global rescheduling. Using synthetic task graphs alongside RIoTBench and WfCommons benchmarks, we systematically evaluate the cost-benefit trade-offs of varying repair scopes across diverse communication-to-computation ratios and interference patterns. Our results demonstrate that no single optimal strategy exists; instead, dynamic adaptation to specific interference characteristics is essential. Rerouting recovers 30% of losses from isolated faults, while global repair approaches optimality under severe congestion. Furthermore, adaptive strategies significantly outperform static baselines.