arXiv · 2610.11260
LLM Agents as Resilience Engineers for Scientific Applications
Abstract
Efficient checkpoint/restart support is essential for resilient HPC scientific applications, but implementing it requires substantial expertise: developers must identify recoverable state, choose globally consistent checkpoint points, and preserve application invariants during restart. We study whether frontier LLM coding agents can automate this process. We build a benchmark suite of 16 MPI applications spanning diverse domains, code sizes, and critical-state structures, and evaluate them with a no-human-in-the-loop generate--validate--revise pipeline for checkpoint/restart synthesis. Across the benchmark, the pipeline produces 41 working resilient implementations. Our results show that agent-driven resilience engineering is practical when critical state is visible or accessible through coherent abstractions: successful runs finish in under one hour on average, consume about 15M tokens, and produce implementations with negligible failure-free overhead and recovery efficiency comparable to human-written code. However, modularized and fragmented state remains a major limitation, with some failed attempts consuming over 100M tokens and 300 minutes without producing a working implementation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hai Duc Nguyen, Tekin Bicer, Kyle Chard, Ian Foster, Bogdan Nicolae. 2026-10-08. LLM Agents as Resilience Engineers for Scientific Applications. https://arxiv.org/abs/2610.11260
Cite the original work for its findings. Save a collection to share your selection of sources.