LLM Agents as Resilience Engineers for Scientific Applications

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high implementation barrier and reliance on domain expertise associated with checkpoint/restart mechanisms in HPC scientific applications by proposing an automated fault-tolerance framework driven by large language model (LLM) coding agents. The framework establishes a generate-verify-revise closed-loop pipeline that autonomously injects fault-tolerance capabilities into MPI applications without human intervention. This work provides the first demonstration that LLM agents can efficiently perform resilience engineering given state visibility. Experimental results show that the proposed approach successfully produces 41 functional implementations, each requiring under one hour on average, while introducing negligible runtime overhead. Furthermore, the achieved fault recovery efficiency is comparable to that of manually written code, substantially reducing development costs.
📝 Abstract
Efficient checkpoint/restart support is essential for resilient HPC scientific applications, but implementing it requires substantial expertise: developers must identify recoverable state, choose globally consistent checkpoint points, and preserve application invariants during restart. We study whether frontier LLM coding agents can automate this process. We build a benchmark suite of 16 MPI applications spanning diverse domains, code sizes, and critical-state structures, and evaluate them with a no-human-in-the-loop generate--validate--revise pipeline for checkpoint/restart synthesis. Across the benchmark, the pipeline produces 41 working resilient implementations. Our results show that agent-driven resilience engineering is practical when critical state is visible or accessible through coherent abstractions: successful runs finish in under one hour on average, consume about 15M tokens, and produce implementations with negligible failure-free overhead and recovery efficiency comparable to human-written code. However, modularized and fragmented state remains a major limitation, with some failed attempts consuming over 100M tokens and 300 minutes without producing a working implementation.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
resilience engineering
checkpoint/restart
high-performance computing
MPI applications
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Agents
Checkpoint/Restart
Resilience Engineering
HPC
Generate-Validate-Revise Pipeline
🔎 Similar Papers