Protecting Futures against Silent Data Corruption -- Efficient Task Replication for Dynamic Data Dependencies

📅 2026-06-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses silent data corruption (SDC) in highly dynamic asynchronous multitasking (AMT) environments characterized by runtime task generation and evolving task dependencies. It proposes a tightly coupled redundancy mechanism that integrates primary and replica computations within an asynchronous task runtime based on C++11 futures/promises and work-stealing load balancing. The approach performs runtime cross-validation of all output effects and selectively re-executes only the affected tasks upon SDC detection. As the first solution to enable efficient SDC protection in such dynamic AMT settings, it supports conditional task spawning and dynamic task graphs while incurring less than 2× performance overhead in fault-free execution—partly due to improved load balancing—and achieves SDC recovery with an overhead of approximately 0.5% of total execution time per incident.
📝 Abstract
As the size of computational problems grows, so does the likelihood of Silent Data Corruptions (SDCs). A common defense is replication, where the computation is repeated and correct results are determined by majority voting. Asynchronous Many-Task (AMT) runtimes are generally well suited for this approach, since the inputs and outputs of task replicas can be compared, and the tasks can be recomputed if necessary. Most existing SDC protection schemes assume static tasks and dependencies. Dynamic settings are more challenging, especially in clusters, since the tasks/data must be tracked for the comparisons. This paper considers a particularly dynamic setting with task spawning at runtime, task communication through C++11-like promises/futures, conditional touches, and cluster-wide load balancing via work-first work stealing. We propose an approach that closely couples original and replica computations by cross-validating all outgoing effects when interacting with the runtime system. The approach selectively recomputes affected tasks only. We implemented the approach in the ItoyoriFBC runtime system and conducted preliminary experiments with Fibonacci and emulated $\mathcal{H}$-matrix LU decomposition benchmarks. Results show a factor of less than two increase of failure-free running times, despite full replication, which is mainly due to improved opportunities for load balancing resulting from the higher number of tasks. The overhead for failure correction was about 0.5% of the overall running time per SDC.
Problem

Research questions and friction points this paper is trying to address.

Silent Data Corruption
Dynamic Task Dependencies
Task Replication
Asynchronous Many-Task Runtime
Fault Tolerance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Silent Data Corruption
Task Replication
Dynamic Task Dependencies
Asynchronous Many-Task Runtime
Selective Recomputation
💼 Related Jobs
No related jobs found.
R
Rüdiger Nather
Research Group Programming Languages / Methodologies, University of Kassel, Kassel, Germany
C
Claudia Fohry
Research Group Programming Languages / Methodologies, University of Kassel, Kassel, Germany
M
Mia Reitz
Research Group Programming Languages / Methodologies, University of Kassel, Kassel, Germany