Silent Data Corruption Protection through Efficient Task Replication

📅 2026-05-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the heightened risk of silent data corruption (SDC) in large-scale supercomputing clusters, where existing replication-based fault-tolerance mechanisms struggle to accommodate asynchronous many-task (AMT) runtimes that support dynamic task generation and work stealing. The authors propose a lightweight SDC detection and recovery mechanism tailored for nested fork-join programs. By recording the task dependency tree, comparing results, and performing a top-down identification of corrupted tasks, the approach selectively re-executes only the affected tasks while reusing correct results from their subtasks. This method achieves precise, localized recovery under dynamic scheduling for the first time, substantially reducing fault-tolerance overhead. Experimental results demonstrate negligible detection and recovery costs, correctness guarantees, and extensibility to future-based task models.
📝 Abstract
The trend of increasing cluster sizes of supercomputers leads to a growing susceptibility to Silent Data Corruption (SDC) that can invalidate program results. A common strategy for SDC protection is replication, where the computation is repeated, and the correct result is determined as the one that is the same in at least two different computations. Applying replication to Asynchronous Many-Task (AMT) runtimes on clusters is challenging due to dynamic task spawning and work stealing, which complicate the identification of replicated tasks. To address the challenge, this paper introduces a novel replication scheme that detects and corrects SDCs for nested fork-join programs. Briefly stated, our approach replicates the computation and records the task tree. Upon a mismatch in the final result, it traverses the tree top-down to identify all corrupted tasks that could have impacted the final result. Recovery is then performed by recomputing these tasks, while the results of correct child tasks are reused. We demonstrate our implementation within a variant of the Itoyori cluster AMT runtime. Our experimental results suggest that the time to identify and reprocess the affected tasks is negligible. The paper concludes by discussing the adaptability of our scheme to tasks that cooperate through futures.
Problem

Research questions and friction points this paper is trying to address.

Silent Data Corruption
Task Replication
Asynchronous Many-Task
Fault Tolerance
Supercomputing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Silent Data Corruption
Task Replication
Asynchronous Many-Task
Task Tree
Fault Recovery
M
Mia Reitz
Research Group Programming Languages / Methodologies, University of Kassel, Kassel, Germany
C
Claudia Fohry
Research Group Programming Languages / Methodologies, University of Kassel, Kassel, Germany