arXiv ScienceSearch

arXiv · 2408.10109

On the loss of orthogonality in low-synchronization variants of reorthogonalized block classical Gram-Schmidt

Abstract

Interest in communication-avoiding orthogonalization schemes for high-performance computing has been growing recently. This manuscript addresses open questions about the numerical stability of various block classical Gram-Schmidt variants that have been proposed in the past few years. An abstract framework is employed, the flexibility of which allows for new rigorous bounds on the loss of orthogonality in these variants. We first analyze a generalization of (reorthogonalized) block classical Gram-Schmidt and show that a "strong" intrablock orthogonalization routine is only needed for the very first block in order to maintain orthogonality on the level of the unit roundoff. In particular, this ``strong" first step does not have to be a reorthogonalized QR itself and subsequent steps can use less stable QR variants, thus keeping the overall communication costs low. Then, using this variant, which has four synchronization points per block column, we remove the synchronization points one at a time and analyze how each alteration affects the stability of the resulting method. Our analysis shows that the variant requiring only one synchronization per block column cannot be guaranteed to be stable in practice, as stability begins to degrade with the first reduction of synchronization points. Our analysis of block methods also provides new theoretical results for the single-column case. In particular, it is proven that DCGS2 from [Bielich, D. et al. Par. Comput. 112 (2022)] and CGS-2 from [Świrydowicz, K. et al, Num. Lin. Alg. Appl. 28 (2021)] are as stable as Householder QR. Numerical examples from the BlockStab toolbox are included throughout, to help compare variants and illustrate the effects of different choices of intraorthogonalization subroutines.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Erin Carson, Kathryn Lund, Yuxin Ma, Eda Oktay. 2025-12-05. On the loss of orthogonality in low-synchronization variants of reorthogonalized block classical Gram-Schmidt. https://arxiv.org/abs/2408.10109

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A Stabilized Finite Element Method for a Morpho-Visco-Poroelastic Model

Studying the structure of soft tissues is important and relevant in biology, particularly in some diseases, such as tumor growth and dermal contraction after burn injury. Based on the complicated characteristics of the tissue and for the sake of a better understanding of the underlying biomechanics, we propose a mathematical model that combines elastic, viscous, and porous effects with growth or shrinkage due to microstructural changes. The framework is referred to as morpho-visco-poroelasticity. Although the existence results of the solution to the problem are not given in this study, we assess the stability of the equilibria for both the continuous and semi-discrete versions of the model, and the key features of this modelling framework have been discussed. To obtain reliable numerical solutions, a stabilized finite element (FE) scheme is proposed for the morpho-visco-poroelasticity equations to avoid spurious oscillations in the pressure profile; the success of this FE scheme is verified by numerical simulations and convergence investigation in both spatial and temporal aspects. For a more quantitative assessment, the total variation of the pressure profile is evaluated as a function of the stabilization parameter.

math.NA

Efficient third-order iterative algorithms for computing zeros of special functions

This manuscript presents a novel and reliable third-order iterative procedure for computing the zeros of solutions to second-order ordinary differential equations. By approximating the solution of the related Riccati differential equation using the trapezoidal rule, this study has derived the proposed third-order method. This work establishes sufficient conditions to ensure the theoretical non-local convergence of the proposed method. This study provides suitable initial guesses for the proposed third-order iterative procedure to compute all zeros in a given interval of the solutions to second-order ordinary differential equations. The orthogonal polynomials like Legendre and Hermite, as well as the special functions like Bessel, Coulomb wave, confluent hypergeometric, and cylinder functions, satisfy the proposed conditions for convergence. Numerical simulations demonstrate the effectiveness of the proposed theory. This work also presents a comparative analysis with recent studies.

math.NA

Machine-Learning-Enhanced Discretize-then-Project Reduced-Order Modeling of Turbulent Flows on Collocated Grids

This study presents a hybrid reduced-order modeling (ROM) framework for incompressible flows on collocated finite-volume grids, combining a discretize-then-project consistent-flux formulation for velocity and pressure with a non-intrusive neural-network closure for turbulent viscosity. The intrusive formulation preserves discrete mass conservation and pressure-velocity coupling, while a reduced pressure reference-cell constraint fixes pressure gauge ambiguity. We evaluate Multilayer Perceptron (MLP), Transformer, and Long Short-Term Memory (LSTM) closures. For a three-dimensional lid-driven cavity at $Re=100$, the LSTM-based ROM achieves relative errors of 0.7% in velocity and 4% in turbulent viscosity. At $Re=3200$, a mode-sensitivity study identifies $N=15$ POD modes as the best overall configuration, balancing accuracy, dimension, robustness, and cost. It yields a final relative velocity error of approximately 12.3% and an online wall-clock speedup of approximately $50\times$ over the full-order model; energy and enstrophy errors remain below 11% for all three architectures. This regime requires case-specific neural-network retraining and pressure reference-cell parameter retuning. In a time-extrapolation test trained on $t\in[0,3]$,s and rolled out to $t=6$,s, the ROM remains bounded, although velocity and pressure errors increase beyond the training window. The LSTM turbulent-viscosity closure remains robust, identifying long-horizon pressure accuracy as the main limitation. These results demonstrate the potential of consistent projection-based modeling combined with data-driven turbulence closure for efficient reduced-order simulation.

math.NA