arXiv · 1809.05805
Low synchronization GMRES algorithms
Abstract
Communication-avoiding and pipelined variants of Krylov solvers are critical for the scalability of linear system solvers on future exascale architectures. We present low synchronization variants of iterated classical (CGS) and modified Gram-Schmidt (MGS) algorithms that require one and two global reduction communication steps. Derivations of low synchronization iterated CGS algorithms are based on previous work by Ruhe. Our main contribution is to introduce a backward normalization lag into the compact $WY$ form of MGS resulting in a ${\cal O}(\eps)\kappa(A)$ stable GMRES algorithm that requires only one global synchronization per iteration. The reduction operations are overlapped with computations and pipelined to optimize performance. Further improvements in performance are achieved by accelerating GMRES BLAS-2 operations on GPUs.
Explore related subjects
Keep this discovery
Kasia Swirydowicz, Julien Langou, Shreyas Ananthan, Ulrike Yang, Stephen Thomas. 2018-09-16. Low synchronization GMRES algorithms. https://arxiv.org/abs/1809.05805
Cite the original work for its findings. Save a collection to share your selection of sources.