Schedule Repair for DAG Workflows under Link Disruptions

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses scheduling failures in DAG workflows for IoT systems under adversarial conditions, where link disruptions compromise execution. We propose a multi-level repair strategy spectrum accounting for decision latency, spanning from passive waiting to global rescheduling. Using synthetic task graphs alongside RIoTBench and WfCommons benchmarks, we systematically evaluate the cost-benefit trade-offs of varying repair scopes across diverse communication-to-computation ratios and interference patterns. Our results demonstrate that no single optimal strategy exists; instead, dynamic adaptation to specific interference characteristics is essential. Rerouting recovers 30% of losses from isolated faults, while global repair approaches optimality under severe congestion. Furthermore, adaptive strategies significantly outperform static baselines.
📝 Abstract
Schedules for directed acyclic graph (DAG) workflows in networked IoT systems are typically computed assuming a static or generally stable network. In contested and adversarial environments, this assumption is not valid. Links degrade and fail due to mobility, interference, and jamming. We study schedule repair: when a link disruption invalidates part of a schedule, how much of it should be rescheduled? We introduce a spectrum of repair policies that vary in repair scope, how much of the pending schedule each may move: wait out the disruption, reroute data around it, reschedule only the affected tasks locally, or reschedule all pending tasks globally. We evaluate each against an oracle and charge every repair a decision latency proportional to the extent to which it moves. Across 100 workload instances spanning synthetic task graphs, RIoTBench pipelines, and WfCommons scientific workflows, each run at five communication-to-computation ratios (CCRs) and disrupted by processes with deliberately different correlation structure, we find that no single scope wins: rerouting nearly erases isolated failures that cost waiting 30%, global repair comes within 4% of the oracle under jamming blackouts, waiting is favored under memoryless link flapping for larger and communication-heavy workloads (the scheduling analog of route-flap damping), self-healing mobility outages reward patience over reaction, and accounting for repair latency erodes large scopes first. We conclude that the scope of the repair should be adapted to the disruption process and the repair cost, rather than fixed by the scheduler.
Problem

Research questions and friction points this paper is trying to address.

Schedule Repair
DAG Workflows
Link Disruptions
IoT Systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Schedule Repair
DAG Workflows
Link Disruptions
Repair Scope
Decision Latency
🔎 Similar Papers
No similar papers found.